Benchmyrk

Guide

How to read an A/B test result

A results page answers one question: is the variant better than what you have, and how sure can you be? Here is what each number means, in the order worth reading them.

Benchmyrk teamPublished 7 min read

Start with the verdict

Every goal in a test gets one of five verdicts. The verdict already combines the statistics below with a few rules about how much data is enough, so read it first and use the numbers to understand why.

The five verdicts Benchmyrk shows for a goal
VerdictWhat it meansWhat to do
Not enough data yetFewer than 7 full days, fewer than 100 visitors in a variant, or fewer than 5 converting visitors in every variant.Wait. Early numbers swing a lot.
Still collecting evidenceEnough data to look at, but no clear answer yet, or the signals disagree.Keep it running until the planned sample size.
Winner foundA variant beats the control on this goal, and the frequentist and Bayesian views agree.Check the traffic split and your guardrails, then ship it.
Variant underperformingEvery variant does worse than the control, with the same level of evidence.Stop the test and keep the control.
No meaningful differenceThe planned sample size is reached and nothing is significant.Keep the simpler version and record what you learned.

A test usually has one primary goal (often orders or revenue per visitor) and a few secondary ones. Decide on the primary goal before launch and judge the test on it. Secondary goals explain the result; they don't overrule it.

Conversion rate and lift

The conversion rate is the share of visitors in a variant who converted on the goal at least once: converting visitors divided by visitors. Benchmyrk counts visitors, not page views, so one shopper who reloads the page ten times counts once.

The lift is the variant's rate relative to the control's. A control at 3.00% and a variant at 3.50% is a lift of +16.6%. Lift is always relative: a "+17%" lift on a 3.0% rate is about half a percentage point.

The 95% interval: how big the lift could be

The observed lift is one estimate. The 95% interval next to it is the range of lifts that fit the data. A narrow interval means the estimate is precise; a wide one means it could still move a lot.

  • If the whole interval is above zero, the data rule out "no effect" at the 95% level for this comparison.
  • If it includes zero, the variant could be better, the same or worse.
  • Its width shrinks roughly with the square root of the number of visitors: four times the traffic halves it.

Read the low end as well as the middle. A lift of +16% with an interval from +2% to +34% says the change helps, but it could be a small help. Plan the rollout around the low end, not the headline number.

Chance to beat control

The chance to beat control is a Bayesian probability: given the data so far, how likely the variant's true rate is higher than the control's. Benchmyrk calculates it exactly from a Beta-Binomial model for conversion goals.

It is not the chance that the lift is as large as observed, and it says nothing about how big the difference is. A variant can have a 99% chance to beat control with a lift that is too small to matter. That is why it sits next to the interval, and why a winner needs both: at least a 95% chance to beat control and a significant p-value.

Why a significant p-value isn't a verdict yet

A p-value is the probability of seeing a difference at least this large if the variant truly made no difference. The classic rule is to call a result when p is below 0.05. That rule assumes you look once, at a sample size you fixed in advance.

Most teams look every day and would stop the moment the number turns green. Looked at that way, the classic p-value finds false winners far more often than 5% of the time (our A/A simulation puts it near one in five over four weeks of daily checks). So until a test reaches its planned sample size, Benchmyrk decides on an always-valid p-value, which stays valid however often anyone looks. After that, the classic fixed-horizon test applies.

The same rates, two amounts of evidenceExample, computed
Control
12,400 visitors, 372 orders (3.00%)
Variant
12,350 visitors, 432 orders (3.50%)
Lift and 95% interval
+16.6% (+1.7% to +33.6%)
Chance to beat control
98.6%
Fixed-horizon p-value
0.027
Always-valid p-value
0.350
Verdict after 14 days
Still collecting evidence

The classic p-value (0.027) and the chance to beat control (98.6%) both look decisive. But the planned sample size for a 3% baseline and a 10% minimum detectable effect is 53,211 visitors per variant, and only 12,350 have arrived. Before that point the always-valid p-value decides, and at 0.350 it isn't there yet.

With the same rates and twice the visitors (24,800 and 24,700), the interval narrows to +5.9% to +28.4%, the always-valid p-value falls to 0.047, the chance to beat control is 99.9%, and the verdict is Winner found.

Example numbers, computed with Benchmyrk's statistics engine at default settings (95% confidence, 80% power, 10% minimum detectable effect). Not data from a store.

With two or more variants, each comparison with the control gets its own p-value, and the more variants you add, the more likely one of them looks good by chance. Benchmyrk adjusts the p-values with the Holm correction, so adding variants doesn't raise the chance of a false winner.

Checks before you trust a result

  • The traffic split. If you planned 50/50 and the variants got noticeably different numbers of visitors, something is filtering visitors unevenly and the result can't be trusted. Benchmyrk runs a sample ratio check and flags the test "Check traffic split" instead of showing a winner. See the SRM guide.
  • Full weeks. Shoppers behave differently on weekdays and weekends. Benchmyrk gives no verdict before 7 full days; running whole weeks is better still.
  • Guardrails. A variant that lifts orders but raises refunds or hurts a guardrail metric isn't a win. Check the guardrail goals before shipping.
  • What changed during the test. A sale, a campaign or a theme update that hits one variant harder than the other can move the numbers. Note them with the decision; the audit log shows changes made to the test itself.

When revenue disagrees with conversion

Conversion rate counts orders; revenue per visitor counts money. They can point different ways: a discount banner can raise the conversion rate and lower revenue per visitor, and a premium bundle can do the opposite. For anything that touches price, offers or order size, judge the test on revenue or gross profit per visitor.

Revenue per visitor is noisier than conversion rate, because order values vary, so it needs more visitors for the same certainty. Benchmyrk compares it with Welch's t-test, which doesn't assume both variants have the same spread, and takes refunds and cancellations off the variant that made the sale. Gross profit per visitor also takes off product costs from Shopify, payment fees and shipping costs; there is no profit verdict while costs are known for less than 90% of order value.

Segments are leads, not results

Breaking a result down by device, new or returning visitors, traffic source or country is useful for ideas. It is not a way to rescue a test. With ten segments, one will usually look like a winner by chance. Benchmyrk labels segment results exploratory and doesn't adjust their intervals for the number of segments. If mobile shoppers seem to love the variant, run a test targeted at mobile to confirm it.

Make the call and write it down

  1. Read the verdict on the primary goal.
  2. Check the traffic split, the days run and the guardrails.
  3. Look at the interval's low end to size the decision.
  4. Ship the winner (a staged rollout lets guardrails watch each step), or keep the control.
  5. Record the decision and why, so the next test starts from it.

A test that ends with no meaningful difference is still an answer: the change didn't matter enough to detect at the size you planned for. Keep whichever version is simpler to maintain, and use the sample size calculator to plan a bolder test next time.

Sources

  1. Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020
  2. Johari, Koomen, Pekelis and Walsh, "Peeking at A/B Tests: Why it matters, and what to do about it", KDD 2017
  3. Holm, "A Simple Sequentially Rejective Multiple Test Procedure", Scandinavian Journal of Statistics, 1979
  4. Welch, "The generalization of 'Student's' problem when several different population variances are involved", Biometrika, 1947

Put it to work on your store.

We'll walk through results and setup on demo data and answer your questions. Or start the 14-day free trial on Shopify.