Benchmyrk

Glossary

Frequentist A/B testing

Analysing a test with p-values and confidence intervals, which describe how surprising the data would be if the variant made no difference.

Frequentist testing asks: if there were no real difference, how often would we see data at least this extreme? That probability is the p-value. A result is significant when it falls below the chosen level, usually 5%.

Its guarantee is about the procedure: it calls false winners at the chosen rate if you plan the sample and don't stop at the first good-looking look, unless the method is built for repeated looks.

For exampleExample

A two-proportion z-test on 25,000 visitors per arm gives p = 0.002 for the new checkout copy.

All terms