Multiple comparisons
The problem that the more comparisons you make in one test, the more likely one looks significant by chance.
Each comparison at 95% confidence has a 5% chance of a false positive. Compare four variants with the control, or slice a result into ten segments, and the chance that at least one looks significant by luck rises well above 5%.
Corrections such as Holm's adjust the p-values so the chance of any false winner stays at the chosen level.
A test with five variants and no correction has about a 23% chance that at least one beats the control by luck alone.
Related terms
- Holm correctionA way to adjust p-values when a test compares several variants with the control, so the chance of any false winner stays at the chosen level.
- False positiveCalling a winner that isn't real: the test finds a difference although the variant made none.
- Exploratory analysisLooking at test results in ways that weren't planned, such as by segment, to find ideas rather than to confirm a result.
- A/B/n testAn A/B test with more than one challenger: the control is compared with two or more variants in the same experiment.