Start with the verdict
Every goal in a test gets one of five verdicts. The verdict already combines the statistics below with a few rules about how much data is enough, so read it first and use the numbers to understand why.
| Verdict | What it means | What to do |
|---|---|---|
| Not enough data yet | Fewer than 7 full days, fewer than 100 visitors in a variant, or fewer than 5 converting visitors in every variant. | Wait. Early numbers swing a lot. |
| Still collecting evidence | Enough data to look at, but no clear answer yet, or the signals disagree. | Keep it running until the planned sample size. |
| Winner found | A variant beats the control on this goal, and the frequentist and Bayesian views agree. | Check the traffic split and your guardrails, then ship it. |
| Variant underperforming | Every variant does worse than the control, with the same level of evidence. | Stop the test and keep the control. |
| No meaningful difference | The planned sample size is reached and nothing is significant. | Keep the simpler version and record what you learned. |
A test usually has one primary goal (often orders or revenue per visitor) and a few secondary ones. Decide on the primary goal before launch and judge the test on it. Secondary goals explain the result; they don't overrule it.
Conversion rate and lift
The conversion rate is the share of visitors in a variant who converted on the goal at least once: converting visitors divided by visitors. Benchmyrk counts visitors, not page views, so one shopper who reloads the page ten times counts once.
The lift is the variant's rate relative to the control's. A control at 3.00% and a variant at 3.50% is a lift of +16.6%. Lift is always relative: a "+17%" lift on a 3.0% rate is about half a percentage point.
The 95% interval: how big the lift could be
The observed lift is one estimate. The 95% interval next to it is the range of lifts that fit the data. A narrow interval means the estimate is precise; a wide one means it could still move a lot.
- If the whole interval is above zero, the data rule out "no effect" at the 95% level for this comparison.
- If it includes zero, the variant could be better, the same or worse.
- Its width shrinks roughly with the square root of the number of visitors: four times the traffic halves it.
Read the low end as well as the middle. A lift of +16% with an interval from +2% to +34% says the change helps, but it could be a small help. Plan the rollout around the low end, not the headline number.
Chance to beat control
The chance to beat control is a Bayesian probability: given the data so far, how likely the variant's true rate is higher than the control's. Benchmyrk calculates it exactly from a Beta-Binomial model for conversion goals.
It is not the chance that the lift is as large as observed, and it says nothing about how big the difference is. A variant can have a 99% chance to beat control with a lift that is too small to matter. That is why it sits next to the interval, and why a winner needs both: at least a 95% chance to beat control and a significant p-value.
Why a significant p-value isn't a verdict yet
A p-value is the probability of seeing a difference at least this large if the variant truly made no difference. The classic rule is to call a result when p is below 0.05. That rule assumes you look once, at a sample size you fixed in advance.
Most teams look every day and would stop the moment the number turns green. Looked at that way, the classic p-value finds false winners far more often than 5% of the time (our A/A simulation puts it near one in five over four weeks of daily checks). So until a test reaches its planned sample size, Benchmyrk decides on an always-valid p-value, which stays valid however often anyone looks. After that, the classic fixed-horizon test applies.
- Control
- 12,400 visitors, 372 orders (3.00%)
- Variant
- 12,350 visitors, 432 orders (3.50%)
- Lift and 95% interval
- +16.6% (+1.7% to +33.6%)
- Chance to beat control
- 98.6%
- Fixed-horizon p-value
- 0.027
- Always-valid p-value
- 0.350
- Verdict after 14 days
- Still collecting evidence
The classic p-value (0.027) and the chance to beat control (98.6%) both look decisive. But the planned sample size for a 3% baseline and a 10% minimum detectable effect is 53,211 visitors per variant, and only 12,350 have arrived. Before that point the always-valid p-value decides, and at 0.350 it isn't there yet.
With the same rates and twice the visitors (24,800 and 24,700), the interval narrows to +5.9% to +28.4%, the always-valid p-value falls to 0.047, the chance to beat control is 99.9%, and the verdict is Winner found.
Example numbers, computed with Benchmyrk's statistics engine at default settings (95% confidence, 80% power, 10% minimum detectable effect). Not data from a store.
With two or more variants, each comparison with the control gets its own p-value, and the more variants you add, the more likely one of them looks good by chance. Benchmyrk adjusts the p-values with the Holm correction, so adding variants doesn't raise the chance of a false winner.
Checks before you trust a result
- The traffic split. If you planned 50/50 and the variants got noticeably different numbers of visitors, something is filtering visitors unevenly and the result can't be trusted. Benchmyrk runs a sample ratio check and flags the test "Check traffic split" instead of showing a winner. See the SRM guide.
- Full weeks. Shoppers behave differently on weekdays and weekends. Benchmyrk gives no verdict before 7 full days; running whole weeks is better still.
- Guardrails. A variant that lifts orders but raises refunds or hurts a guardrail metric isn't a win. Check the guardrail goals before shipping.
- What changed during the test. A sale, a campaign or a theme update that hits one variant harder than the other can move the numbers. Note them with the decision; the audit log shows changes made to the test itself.
When revenue disagrees with conversion
Conversion rate counts orders; revenue per visitor counts money. They can point different ways: a discount banner can raise the conversion rate and lower revenue per visitor, and a premium bundle can do the opposite. For anything that touches price, offers or order size, judge the test on revenue or gross profit per visitor.
Revenue per visitor is noisier than conversion rate, because order values vary, so it needs more visitors for the same certainty. Benchmyrk compares it with Welch's t-test, which doesn't assume both variants have the same spread, and takes refunds and cancellations off the variant that made the sale. Gross profit per visitor also takes off product costs from Shopify, payment fees and shipping costs; there is no profit verdict while costs are known for less than 90% of order value.
Segments are leads, not results
Breaking a result down by device, new or returning visitors, traffic source or country is useful for ideas. It is not a way to rescue a test. With ten segments, one will usually look like a winner by chance. Benchmyrk labels segment results exploratory and doesn't adjust their intervals for the number of segments. If mobile shoppers seem to love the variant, run a test targeted at mobile to confirm it.
Make the call and write it down
- Read the verdict on the primary goal.
- Check the traffic split, the days run and the guardrails.
- Look at the interval's low end to size the decision.
- Ship the winner (a staged rollout lets guardrails watch each step), or keep the control.
- Record the decision and why, so the next test starts from it.
A test that ends with no meaningful difference is still an answer: the change didn't matter enough to detect at the size you planned for. Keep whichever version is simpler to maintain, and use the sample size calculator to plan a bolder test next time.
Sources
- Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020
- Johari, Koomen, Pekelis and Walsh, "Peeking at A/B Tests: Why it matters, and what to do about it", KDD 2017
- Holm, "A Simple Sequentially Rejective Multiple Test Procedure", Scandinavian Journal of Statistics, 1979
- Welch, "The generalization of 'Student's' problem when several different population variances are involved", Biometrika, 1947