Benchmyrk

Guide

Sample size and test duration: how long to run an A/B test

Before a test starts, decide how many visitors it needs. That number tells you how long to run it, whether the test is worth running at all, and when a result is ready.

Benchmyrk teamPublished 5 min read

The four inputs

  • Baseline conversion rate. How often visitors to the tested pages convert today, for the goal you'll judge the test on. Take it from the last few weeks, from the same pages and devices the test will target.
  • Minimum detectable effect (MDE). The smallest lift worth detecting, relative to the baseline. A 10% MDE on a 3% rate means detecting 3.3%. Smaller effects need many more visitors.
  • Significance (confidence level). How strict the test is about false winners. 95% confidence means a 5% chance of calling a winner when there is no difference.
  • Power. The chance of detecting the effect if it is really there at the MDE. 80% is the usual choice: one real effect in five of that size is missed.

Then the number of variants: each challenger is compared with the control, and more comparisons need more visitors per variant to keep false winners rare.

The formula Benchmyrk uses

Benchmyrk plans conversion goals with the standard sample size for a two-sided, pooled two-proportion z-test (Fleiss, Levin and Paik, equation 4.14, without continuity correction). It is the same test the results engine uses once the planned sample is reached, so the plan matches how the result is judged.

n = [ z(1 − α′/2) · √(2·p̄·(1 − p̄)) + z(1 − β) · √(p₁(1 − p₁) + p₂(1 − p₂)) ]² / (p₂ − p₁)²
n: visitors per variant. p₁: baseline rate. p₂ = p₁ · (1 + MDE). p̄ = (p₁ + p₂) / 2. α′ = α / (number of challengers). z: the standard normal quantile. 1 − β: power.

With more than one challenger, the significance level is split evenly between the comparisons (a Bonferroni split): two challengers at 95% confidence plan each comparison at 97.5%. Results are later adjusted with the Holm correction, which is never stricter than this split, so the plan errs on the safe side.

A worked example

Product page test on a store with 4,000 daily visitorsExample, computed
Baseline conversion rate
3.0%
Minimum detectable effect
10% relative (3.0% → 3.3%)
Confidence and power
95% and 80%
Variants
Control and one challenger, 50/50
Visitors per variant
53,211
Visitors in total
106,422
Days at 4,000 visitors a day
27

Computed with Benchmyrk's planning code (the same numbers the new-experiment wizard shows). Example inputs, not data from a store.

Each variant needs 53,211 visitors. With every visitor entering the test and a 50/50 split, each variant gets 2,000 a day, so the test needs 27 days. Round that up to whole weeks: 28 days covers 4 full weekly cycles.

Adding a second challenger doesn't just add a third of the traffic. Each variant now needs 64,439 visitors (the significance level is split between two comparisons), 193,317 in total, and at the same traffic the test takes 49 days.

What moves the number most

The MDE matters most: halving it roughly quadruples the sample, because the required sample grows with one over the effect squared.

Visitors needed at a 3% baseline, 95% confidence and 80% power, by minimum detectable effect
Minimum detectable effectPer variantTotal (two variants)
5%207,938415,876
10%53,211106,422
20%13,91427,828
Computed with Benchmyrk's statistics engine.

A lower baseline needs more visitors for the same relative lift, because conversions are rarer. Checkout and purchase goals usually have the lowest rates; add-to-cart rates are higher, which is one reason teams test on them, but a lift in add-to-cart doesn't always reach orders.

Visitors needed for a 10% relative lift, 95% confidence and 80% power, by baseline rate
Baseline ratePer variantTotal (two variants)
1%163,095326,190
2%80,682161,364
3%53,211106,422
5%31,23462,468
Computed with Benchmyrk's statistics engine.

From visitors to days

days = visitors per variant / (daily visitors × share in the test × share of the smallest variant)
Every variant has to reach the planned size, so the smallest share sets the pace. Round up.
  • Run at least 7 full days, and whole weeks where you can. Benchmyrk gives no verdict before 7 full days.
  • Count only visitors who reach the tested pages. A test on product pages gets product-page visitors, not the whole store's traffic.
  • Avoid starting during an unusual week (a big sale or a launch) unless that is what you want to learn about.
  • If the plan comes out longer than about four weeks, Benchmyrk flags it as low traffic: cookies expire, shoppers switch devices and seasons change over that time.

When the number is too big

  • Test a bigger change. A new offer or a reworked page can move the rate by 20% or more; a button colour rarely does. A larger MDE needs far fewer visitors.
  • Test where the traffic is. A change on every product page or in the site header reaches more visitors than one landing page.
  • Use fewer variants. One strong challenger beats three weak ones on a small store.
  • Reduce variance. For purchase and revenue goals, CUPED uses each visitor's orders before the test to tighten the intervals. It doesn't change the plan, but it can bring a result forward.
  • Use auto-allocation for short campaigns. When the goal is to earn the most during a promotion rather than to learn a precise lift, an auto-allocated test moves traffic toward the leader as it runs.

Peeking and the planned sample

The planned sample is not a lock on the results page. You can look as often as you like: until every variant reaches the planned size, Benchmyrk only calls a winner on an always-valid p-value, which allows for repeated looks. A large effect can be called early that way. A small one usually needs the full sample, and once it is reached the classic test applies. Read more in How to read an A/B test result.

Planning for revenue per visitor

Revenue per visitor has no single rate to plan from: most visitors spend nothing and buyers spend different amounts, so the sample depends on how spread out spending is (its variance). The usual formula for comparing two means applies:

n = 2σ² · (z(1 − α′/2) + z(1 − β))² / Δ²
σ²: variance of revenue per visitor. Δ: the smallest difference worth detecting, MDE × baseline revenue per visitor.

Benchmyrk plans revenue goals with this formula once the test has data, using the control's observed variance. Expect revenue goals to need more visitors than the conversion rate of the same test.

Try your own numbers

The sample size calculator runs the same code in your browser: enter your baseline, MDE, variants and daily visitors, and share the link with your team.

Sources

  1. Fleiss, Levin and Paik, Statistical Methods for Rates and Proportions, 3rd edition, Wiley, 2003
  2. Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020
  3. Deng, Xu, Kohavi and Walker, "Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data", WSDM 2013

Put it to work on your store.

We'll walk through results and setup on demo data and answer your questions. Or start the 14-day free trial on Shopify.