Benchmyrk

Guide

SRM: when your A/B test split looks wrong

You planned a 50/50 split and one variant got noticeably fewer visitors. That gap is a sample ratio mismatch, and it means something is filtering visitors unevenly. Until you know what, the result can't be trusted, however good it looks.

Benchmyrk teamPublished 4 min read

What a sample ratio mismatch is

Random assignment never splits traffic exactly 50/50, but over thousands of visitors the gap should stay small. A sample ratio mismatch (SRM) is a gap too large to be chance. It is one of the most common ways an A/B test goes wrong, and one of the easiest to check.

Why it matters: the visitors missing from one variant are rarely a random sample. If a variant breaks on one browser, the shoppers it loses are those on that browser, and the comparison is no longer like for like. The bias can make a variant look better or worse than it is.

How Benchmyrk checks the split

Each results update runs a chi-square goodness-of-fit test of the visitors in each variant against the split you configured. It is the same test whether the split is 50/50, 90/10 or three ways.

χ² = Σ (observedᵢ − expectedᵢ)² / expectedᵢ,   expectedᵢ = total visitors × planned shareᵢ
Degrees of freedom: number of variants − 1. The p-value is the chance of a gap at least this large if the split were working.

A test is flagged when p is below 0.001 and at least 200 visitors have entered. The threshold is strict on purpose: the check runs again with every update, and a looser one would raise false alarms on healthy tests.

Three 50/50 testsExample, computed
10,000 vs 9,850 visitors
1.5% gap, p 0.287: fine
10,000 vs 9,600 visitors
4.0% gap, p 0.0043: not flagged, worth a look
10,000 vs 9,400 visitors
6.0% gap, p < 0.001: flagged

A 1.5% gap on twenty thousand visitors is ordinary chance (p 0.287). At 4.0% (p 0.0043) it is unusual but not past the threshold: check the daily view to see whether the gap is growing. At 6.0% (p < 0.001) chance is no longer a credible explanation and the test is flagged.

Example counts; p-values computed with Benchmyrk's statistics engine.

Common causes on ecommerce stores

  • The variant breaks for some visitors. A script error, a change that fails on one browser or a layout that hides the content on small screens can stop those visitors from being counted in that variant.
  • Redirects. In a split URL test the variant's visitors make an extra hop; slow pages and visitors who leave during the redirect go missing from one side.
  • Different filtering per variant. Bot filtering, consent banners or ad blockers that treat the variants differently (for example, a variant that changes the consent banner).
  • Changes mid-test. Editing the traffic split, the targeting or the variant's pages while the test runs mixes two experiments in one result.
  • Caching. A page cache or CDN that serves one variant's HTML to visitors assigned to the other.
  • Counting from a later step. Exposure recorded on an action that only some visitors take, and that the change itself affects.

Fabijan and colleagues at Microsoft catalogued these causes across many experiments; their taxonomy (linked below) is the best reference when the usual suspects don't explain a gap.

What to do when a test is flagged

  1. Don't act on the result, in either direction.
  2. Open the daily sample ratio view and find when the gap started. A sudden start points at a change on that day; a steady gap points at the variant itself.
  3. Compare the variants on devices and browsers. Preview each variant on desktop, tablet and phone, and look at the launch check's renders of each device.
  4. Check what changed: the test's history in the audit log, theme updates, apps installed, campaigns started.
  5. Fix the cause, then restart the test so the new result is built on clean data. Continuing the old test mixes broken and fixed traffic.

If you can't find a cause, an A/A test (two identical variants) on the same pages tells you whether the problem is in the setup or in the change you tested.

Sources

  1. Fabijan, Gupchup, Gupta, Omhover, Qin, Vermeer and Dmitriev, "Diagnosing Sample Ratio Mismatch in Online Controlled Experiments", KDD 2019
  2. Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020 (chapter 21)

Put it to work on your store.

We'll walk through results and setup on demo data and answer your questions. Or start the 14-day free trial on Shopify.