Results you can defend.
Every result shows the chance to beat control, a 95% interval and the checks that tell you when not to trust it.
Two views of the same data
A winner is called only when both agree.
Bayesian
The chance each variant beats the control, calculated exactly from a Beta-Binomial model, with the expected loss of choosing each variant.
Frequentist
A two-proportion z-test for conversion rates and Welch's t-test for revenue and profit per visitor, with 95% intervals on the lift.
Look as often as you like
Checking a test every day and stopping at the first significant result can multiply false winners. Until a test reaches its planned sample size, Benchmyrk calls a winner only on an always-valid p-value from a mixture sequential probability ratio test (mSPRT), which stays valid however often anyone looks. After that, the fixed-horizon test applies.
- On by default for every fixed-split test.
- The planned sample size comes from your baseline rate and the smallest lift you want to detect, set before launch.
The checks behind every verdict
Enough data
No verdict before 7 full days, 100 visitors in every variant and 5 converting visitors per variant.
Agreement
A winner needs an adjusted p-value under 5% (at the default 95% confidence level), at least a 95% chance to beat control, and a positive lift.
Several challengers
The p-values are Holm-corrected across variants, so adding variants doesn't raise the chance of a false winner.
Sample ratio mismatch
A chi-square test of visitors per variant against the planned split. A mismatch (p < 0.001, at least 200 visitors) is flagged before anyone trusts the result.
Guardrails
Metrics that must not get worse; when one does, Benchmyrk alerts you or pauses the test.
Profit
No profit verdict while product costs are known for less than 90% of order value.
| Check | The rule |
|---|---|
| Enough data | No verdict before 7 full days, 100 visitors in every variant and 5 converting visitors per variant. |
| Agreement | A winner needs an adjusted p-value under 5% (at the default 95% confidence level), at least a 95% chance to beat control, and a positive lift. |
| Several challengers | The p-values are Holm-corrected across variants, so adding variants doesn't raise the chance of a false winner. |
| Sample ratio mismatch | A chi-square test of visitors per variant against the planned split. A mismatch (p < 0.001, at least 200 visitors) is flagged before anyone trusts the result. |
| Guardrails | Metrics that must not get worse; when one does, Benchmyrk alerts you or pauses the test. |
| Profit | No profit verdict while product costs are known for less than 90% of order value. |
Revenue and profit, measured properly
Revenue per visitor
From paid orders, with Welch's t-test, which doesn't assume equal variances.
Refunds taken off
Refunds and cancellations reduce the revenue credited to the variant that sold it.
Gross profit per visitor
Shopify unit costs saved on each order, minus payment fees and shipping costs.
CUPED
Optional variance reduction from each visitor's orders in the 30 days before the test, for tighter intervals. Chosen before launch.
Segments, marked exploratory
Each variant's lift by device, new or returning visitors, traffic source and country, side by side with its interval. Segment intervals aren't adjusted for the number of segments, so the page says so: a segment difference is a lead for the next test, not a result.
The methods, with their sources
Standard, published methods. Formulas are in the engine's source comments.
Always-valid p-values (mSPRT)
Johari, Koomen, Pekelis and Walsh, "Peeking at A/B Tests", KDD 2017, and "Always Valid Inference", Operations Research, 2022.
CUPED
Deng, Xu, Kohavi and Walker, "Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data", WSDM 2013.
Multiple comparisons
Holm, "A Simple Sequentially Rejective Multiple Test Procedure", Scandinavian Journal of Statistics, 1979.
Auto-allocation decisions
Anytime-valid confidence sequences: Howard, Ramdas, McAuliffe and Sekhon, Annals of Statistics, 2021.
| Method | Where it comes from |
|---|---|
| Always-valid p-values (mSPRT) | Johari, Koomen, Pekelis and Walsh, "Peeking at A/B Tests", KDD 2017, and "Always Valid Inference", Operations Research, 2022. |
| CUPED | Deng, Xu, Kohavi and Walker, "Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data", WSDM 2013. |
| Multiple comparisons | Holm, "A Simple Sequentially Rejective Multiple Test Procedure", Scandinavian Journal of Statistics, 1979. |
| Auto-allocation decisions | Anytime-valid confidence sequences: Howard, Ramdas, McAuliffe and Sekhon, Annals of Statistics, 2021. |
Good to know
- Auto-allocated tests report each variant's chance of being best instead of p-values, because a moving split biases them. See auto-allocation.
- CUPED applies to purchase, revenue and profit goals, and only when enough visitors have order history.
- The AI readout of a result has to agree with these statistics, and tells you to investigate first when the split looks broken.
Questions
Bayesian or frequentist?
Both, from the same data. At the default 95% confidence level, a winner needs an adjusted p-value under 5% and at least a 95% chance to beat control, so the two views have to agree.
Can I check results every day?
Yes. Until the planned sample size is reached, winners are called only on an always-valid p-value, which stays valid however often you look.
What does chance to beat control mean?
The probability, given the data so far, that the variant's true rate is higher than the control's. It comes from the Bayesian model and is shown next to the frequentist interval.
Why hasn't my test called a winner?
It may not have run 7 full days, reached 100 visitors and 5 converting visitors in every variant, or shown a difference both views agree on. A test can also end with no difference once it reaches its planned sample size.
Bring your analyst.
We'll go through the method on demo data and answer the hard questions. Or install the Shopify app and read your own results.