False positive
Calling a winner that isn't real: the test finds a difference although the variant made none.
Also called Type I error
A false positive is declaring a significant result when the change had no effect. At 95% confidence, one look at one comparison has a 5% chance of it.
Checking results repeatedly and stopping early, testing many variants or many segments all raise that chance unless the method corrects for them.
A store tests 20 button colours against the control at 95% confidence without correction. One of them "wins" by chance alone.
Related terms
- P-valueThe probability of seeing a difference at least as large as the observed one if the variant actually made no difference.
- PeekingChecking a test's results while it runs and stopping as soon as they look significant, which inflates false winners with classic p-values.
- Multiple comparisonsThe problem that the more comparisons you make in one test, the more likely one looks significant by chance.
- A/A testA test where both arms show the same page, used to check that the split, tracking and statistics behave before trusting real tests.