Benchmyrk

Blog

Checking a test every day: what it does to false winners

Everyone checks a running test. The question is what happens when you act on what you see. We simulated 1,000 tests where the variant changes nothing and counted how often each way of reading them declared a winner anyway.

Benchmyrk teamPublished 3 min read

The setup

  • 1,000 simulated A/A tests: two arms with the same true conversion rate of 3%, so every "winner" is a false one.
  • 1,000 visitors per arm per day for 28 days, with a binomial number of converting visitors each day.
  • From day 7, three ways of reading each test, checked once a day.
  • A fixed random seed, so the run can be repeated exactly. The code is part of our test suite and fails if these numbers change.
The three stopping rules compared
RuleWhat it does
Daily peekingThe classic two-proportion z-test on the data so far, every day; stop at the first p < 0.05.
One lookThe same test once, on day 28, as the method assumes.
BenchmyrkThe verdict from our real results engine at default settings, every day; stop at the first "Winner found" or "Variant underperforming".

The results

False winners in 1,000 simulated A/A tests
RuleTests that called a winnerShare
Daily peeking18918.9%
One look on day 28535.3%
Benchmyrk, checked daily131.3%
Simulated data, not results from stores. With 1,000 tests, the 95% margins are ±2.4, ±1.4 and ±0.7 percentage points.

Checking every day and stopping at the first significant result called a winner in 18.9% of tests where nothing changed: nearly four times the 5% the method promises. Looking once, at the end, behaved as advertised (5.3%). Benchmyrk's verdict, checked just as often as the peeker, called 1.3%.

Why daily checks inflate false winners

A p-value below 0.05 means a difference this large would show up by chance less than 5% of the time, in one look. Every extra look is another draw. Over 22 daily looks, the running result wanders above and below the line, and the chance that it crosses it at least once is far higher than 5%. Stopping at the first crossing turns that wandering into a decision.

How Benchmyrk stays at or under 5%

Until a test reaches its planned sample size, Benchmyrk decides on an always-valid p-value from a mixture sequential probability ratio test (Johari and colleagues, 2017 and 2022). Its guarantee holds however many times anyone looks: the chance of ever crossing 5% in a test with no real difference is at most 5% (under the same large-sample approximation the classic test uses). On top of that, a winner needs at least a 95% chance to beat control and 7 full days of data. In this simulation no test reached its planned sample size (53,211 visitors per arm for a 3% rate and a 10% minimum detectable effect), so every verdict came from the always-valid rule.

The 1.3% is lower than 5% because the guarantee is a bound, not a target: the method is conservative.

The trade-off

Nothing is free. An always-valid test needs stronger evidence than a single fixed look to call the same effect early, so a real but modest lift often waits for the planned sample, where the classic test takes over. In exchange you can look every day, and stop early with confidence when an effect is large. For most store teams, who will look anyway, that is the better deal.

If you'd rather see the rule applied to numbers, How to read an A/B test result walks through a readout where the classic p-value says "significant" and the verdict says "still collecting evidence", and why.

Sources

  1. Johari, Koomen, Pekelis and Walsh, "Peeking at A/B Tests: Why it matters, and what to do about it", KDD 2017
  2. Johari, Koomen, Pekelis and Walsh, "Always Valid Inference: Continuous Monitoring of A/B Tests", Operations Research, 2022

See it on your store.

We'll walk through results and setup on demo data and answer your questions. Or start the 14-day free trial on Shopify.