Benchmyrk

Free tool

A/B test sample size calculator.

How many visitors your test needs, and how many days that takes. The same formula as the Benchmyrk app, calculated in your browser as you type.

%

Today's rate on the tested pages, for your goal.

%

Smallest relative lift worth detecting.

Including the control.

Visitors a day who reach the tested pages.

%

100% unless you hold some visitors back.

Your plan

Visitors per variant
53,211
Visitors in total
106,422
2 variants, control included
Days to run
—
Add daily visitors to see the duration

Powered to detect 3% → 3.3% with 95% confidence and 80% power.

How it's calculated

The standard sample size for a two-sided, pooled two-proportion z-test (Fleiss, Levin and Paik, 2003, equation 4.14), the test Benchmyrk's results use once the planned sample is reached.

n = [ z(1 − α′/2) · √(2·p̄·(1 − p̄)) + z(1 − β) · √(p₁(1 − p₁) + p₂(1 − p₂)) ]² / (p₂ − p₁)²
n: visitors per variant. p₁: baseline rate. p₂ = p₁ · (1 + MDE). p̄ = (p₁ + p₂) / 2. α′ = (1 − confidence) / (variants − 1). 1 − β: power. Rounded up.

Days = visitors per variant ÷ (daily visitors × share in the test ÷ variants), rounded up: every variant has to reach the planned size, so with an even split each gets an equal share of the visitors entering the test.

What it assumes

Questions

Is this the same calculation as the Benchmyrk app?

Yes. The calculator runs the planning code the new-experiment wizard uses, so for the same inputs it gives the same visitors per variant and days.

Can I stop the test before it reaches this number?

With Benchmyrk, yes, if the evidence is strong: until a test reaches its planned sample, winners are called on an always-valid p-value that allows looking every day. Small effects usually need the full sample. With a classic fixed-horizon test, stopping early inflates false winners.

Why does a smaller minimum detectable effect need so many more visitors?

The sample grows with one over the effect squared: halving the effect you want to detect roughly quadruples the visitors needed.

Why run whole weeks?

Shoppers behave differently on weekdays and weekends. Ending on a full week keeps each day of the week equally represented. Benchmyrk gives no verdict before 7 full days.

What about revenue per visitor?

Revenue needs the variance of spending per visitor, which this calculator doesn't ask for. Benchmyrk plans revenue goals from the variance it sees in the running test; expect them to need more visitors than conversion rate.

Plan it, then run it.

Benchmyrk plans every test with this formula and tells you when the result is ready. Book a demo, or start the free trial on Shopify.