The four inputs
- Baseline conversion rate. How often visitors to the tested pages convert today, for the goal you'll judge the test on. Take it from the last few weeks, from the same pages and devices the test will target.
- Minimum detectable effect (MDE). The smallest lift worth detecting, relative to the baseline. A 10% MDE on a 3% rate means detecting 3.3%. Smaller effects need many more visitors.
- Significance (confidence level). How strict the test is about false winners. 95% confidence means a 5% chance of calling a winner when there is no difference.
- Power. The chance of detecting the effect if it is really there at the MDE. 80% is the usual choice: one real effect in five of that size is missed.
Then the number of variants: each challenger is compared with the control, and more comparisons need more visitors per variant to keep false winners rare.
The formula Benchmyrk uses
Benchmyrk plans conversion goals with the standard sample size for a two-sided, pooled two-proportion z-test (Fleiss, Levin and Paik, equation 4.14, without continuity correction). It is the same test the results engine uses once the planned sample is reached, so the plan matches how the result is judged.
n = [ z(1 − α′/2) · √(2·p̄·(1 − p̄)) + z(1 − β) · √(p₁(1 − p₁) + p₂(1 − p₂)) ]² / (p₂ − p₁)²With more than one challenger, the significance level is split evenly between the comparisons (a Bonferroni split): two challengers at 95% confidence plan each comparison at 97.5%. Results are later adjusted with the Holm correction, which is never stricter than this split, so the plan errs on the safe side.
A worked example
- Baseline conversion rate
- 3.0%
- Minimum detectable effect
- 10% relative (3.0% → 3.3%)
- Confidence and power
- 95% and 80%
- Variants
- Control and one challenger, 50/50
- Visitors per variant
- 53,211
- Visitors in total
- 106,422
- Days at 4,000 visitors a day
- 27
Computed with Benchmyrk's planning code (the same numbers the new-experiment wizard shows). Example inputs, not data from a store.
Each variant needs 53,211 visitors. With every visitor entering the test and a 50/50 split, each variant gets 2,000 a day, so the test needs 27 days. Round that up to whole weeks: 28 days covers 4 full weekly cycles.
Adding a second challenger doesn't just add a third of the traffic. Each variant now needs 64,439 visitors (the significance level is split between two comparisons), 193,317 in total, and at the same traffic the test takes 49 days.
What moves the number most
The MDE matters most: halving it roughly quadruples the sample, because the required sample grows with one over the effect squared.
| Minimum detectable effect | Per variant | Total (two variants) |
|---|---|---|
| 5% | 207,938 | 415,876 |
| 10% | 53,211 | 106,422 |
| 20% | 13,914 | 27,828 |
A lower baseline needs more visitors for the same relative lift, because conversions are rarer. Checkout and purchase goals usually have the lowest rates; add-to-cart rates are higher, which is one reason teams test on them, but a lift in add-to-cart doesn't always reach orders.
| Baseline rate | Per variant | Total (two variants) |
|---|---|---|
| 1% | 163,095 | 326,190 |
| 2% | 80,682 | 161,364 |
| 3% | 53,211 | 106,422 |
| 5% | 31,234 | 62,468 |
From visitors to days
days = visitors per variant / (daily visitors × share in the test × share of the smallest variant)- Run at least 7 full days, and whole weeks where you can. Benchmyrk gives no verdict before 7 full days.
- Count only visitors who reach the tested pages. A test on product pages gets product-page visitors, not the whole store's traffic.
- Avoid starting during an unusual week (a big sale or a launch) unless that is what you want to learn about.
- If the plan comes out longer than about four weeks, Benchmyrk flags it as low traffic: cookies expire, shoppers switch devices and seasons change over that time.
When the number is too big
- Test a bigger change. A new offer or a reworked page can move the rate by 20% or more; a button colour rarely does. A larger MDE needs far fewer visitors.
- Test where the traffic is. A change on every product page or in the site header reaches more visitors than one landing page.
- Use fewer variants. One strong challenger beats three weak ones on a small store.
- Reduce variance. For purchase and revenue goals, CUPED uses each visitor's orders before the test to tighten the intervals. It doesn't change the plan, but it can bring a result forward.
- Use auto-allocation for short campaigns. When the goal is to earn the most during a promotion rather than to learn a precise lift, an auto-allocated test moves traffic toward the leader as it runs.
Peeking and the planned sample
The planned sample is not a lock on the results page. You can look as often as you like: until every variant reaches the planned size, Benchmyrk only calls a winner on an always-valid p-value, which allows for repeated looks. A large effect can be called early that way. A small one usually needs the full sample, and once it is reached the classic test applies. Read more in How to read an A/B test result.
Planning for revenue per visitor
Revenue per visitor has no single rate to plan from: most visitors spend nothing and buyers spend different amounts, so the sample depends on how spread out spending is (its variance). The usual formula for comparing two means applies:
n = 2σ² · (z(1 − α′/2) + z(1 − β))² / Δ²Benchmyrk plans revenue goals with this formula once the test has data, using the control's observed variance. Expect revenue goals to need more visitors than the conversion rate of the same test.
Try your own numbers
The sample size calculator runs the same code in your browser: enter your baseline, MDE, variants and daily visitors, and share the link with your team.
Sources
- Fleiss, Levin and Paik, Statistical Methods for Rates and Proportions, 3rd edition, Wiley, 2003
- Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020
- Deng, Xu, Kohavi and Walker, "Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data", WSDM 2013