Peeking
Checking a test's results while it runs and stopping as soon as they look significant, which inflates false winners with classic p-values.
Every look at a classic p-value is another chance for noise to cross the significance line. Stopping at the first crossing turns that noise into decisions: with daily checks over a few weeks, false winners can become several times more common than the 5% the method promises.
The fix is either to look only once, at the planned sample size, or to use a method designed for repeated looks.
A team checks every morning and stops the day p dips below 0.05. In tests where nothing changed, they would declare a winner far more often than one time in twenty.
Related terms
- Sequential testingStatistical methods that let you check a test's results as often as you like and stop early, without inflating false winners.
- False positiveCalling a winner that isn't real: the test finds a difference although the variant made none.
- P-valueThe probability of seeing a difference at least as large as the observed one if the variant actually made no difference.
- Sample sizeHow many visitors a test needs per variant to detect the minimum detectable effect with the chosen confidence and power.