Benchmyrk

Glossary

Peeking

Checking a test's results while it runs and stopping as soon as they look significant, which inflates false winners with classic p-values.

Every look at a classic p-value is another chance for noise to cross the significance line. Stopping at the first crossing turns that noise into decisions: with daily checks over a few weeks, false winners can become several times more common than the 5% the method promises.

The fix is either to look only once, at the planned sample size, or to use a method designed for repeated looks.

For exampleExample

A team checks every morning and stops the day p dips below 0.05. In tests where nothing changed, they would declare a winner far more often than one time in twenty.

Where it comes up

All terms