Sample size
How many visitors a test needs per variant to detect the minimum detectable effect with the chosen confidence and power.
The sample size follows from four inputs: the baseline rate, the minimum detectable effect, the confidence level and the power. Planning it before launch tells you how long the test will take and whether it is worth running.
Ending a test well before its planned sample, without a method built for early stopping, makes false winners and exaggerated lifts more likely.
At a 3% baseline, a 10% MDE, 95% confidence and 80% power, each variant needs 53,211 visitors.
Related terms
- Minimum detectable effect (MDE)The smallest lift a test is planned to detect reliably; it sets how many visitors the test needs.
- Statistical powerThe chance that a test detects an effect of a given size when that effect is really there.
- Statistical significanceA result is significant when its p-value is below a chosen threshold, such as 0.05, meaning chance alone is an unlikely explanation.
- PeekingChecking a test's results while it runs and stopping as soon as they look significant, which inflates false winners with classic p-values.