Benchmyrk

Glossary

Multiple comparisons

The problem that the more comparisons you make in one test, the more likely one looks significant by chance.

Each comparison at 95% confidence has a 5% chance of a false positive. Compare four variants with the control, or slice a result into ten segments, and the chance that at least one looks significant by luck rises well above 5%.

Corrections such as Holm's adjust the p-values so the chance of any false winner stays at the chosen level.

For exampleExample

A test with five variants and no correction has about a 23% chance that at least one beats the control by luck alone.

All terms