Sequential testing
Statistical methods that let you check a test's results as often as you like and stop early, without inflating false winners.
Also called always-valid p-value, continuous monitoring
Classic tests assume one look at a fixed sample. Sequential methods are built for repeated looks: their error guarantee holds at every look at once, so stopping when the evidence is strong is safe.
One family produces always-valid p-values, for example from a mixture sequential probability ratio test (mSPRT). The trade-off is that a modest effect needs somewhat more evidence to be called early.
A team checks a test daily. On day 12 the always-valid p-value drops below 0.05 and they stop with a winner; the guarantee covers all twelve looks.
Related terms
- PeekingChecking a test's results while it runs and stopping as soon as they look significant, which inflates false winners with classic p-values.
- P-valueThe probability of seeing a difference at least as large as the observed one if the variant actually made no difference.
- False positiveCalling a winner that isn't real: the test finds a difference although the variant made none.
- Sample sizeHow many visitors a test needs per variant to detect the minimum detectable effect with the chosen confidence and power.