Frequentist A/B testing
Analysing a test with p-values and confidence intervals, which describe how surprising the data would be if the variant made no difference.
Frequentist testing asks: if there were no real difference, how often would we see data at least this extreme? That probability is the p-value. A result is significant when it falls below the chosen level, usually 5%.
Its guarantee is about the procedure: it calls false winners at the chosen rate if you plan the sample and don't stop at the first good-looking look, unless the method is built for repeated looks.
A two-proportion z-test on 25,000 visitors per arm gives p = 0.002 for the new checkout copy.
Related terms
- P-valueThe probability of seeing a difference at least as large as the observed one if the variant actually made no difference.
- Confidence intervalA range of values for the true effect that is consistent with the data, at a stated confidence level such as 95%.
- Bayesian A/B testingAnalysing a test by updating a probability model with the data, so results read as probabilities such as the chance a variant beats the control.
- Sequential testingStatistical methods that let you check a test's results as often as you like and stop early, without inflating false winners.