The short answer
- At least 7 full days, in whole weeks. Shoppers buy differently on weekdays and weekends; a test that ends mid-week over-represents some days.
- Until each version reaches the planned sample. The plan comes from your conversion rate, the smallest change worth detecting and your traffic. Work it out before you start.
- Rarely more than four weeks. Past about 28 days, cookies get cleared, shoppers switch devices and the season moves on. If the plan is longer, change the test, not the deadline.
- Earlier only on evidence built for repeated looks. Stopping the first day a dashboard shows 95% inflates false winners. A sequential (always-valid) test lets you stop early safely when the effect is large.
The plan sets the length
How long a test runs follows from how many visitors it needs. That number depends mostly on the smallest change you want to be able to detect (the minimum detectable effect): halving it roughly quadruples the visitors.
| Smallest lift worth detecting | Visitors per version | Days needed | Run for |
|---|---|---|---|
| 10% | 53,211 | 36 | 6 weeks |
| 20% | 13,914 | 10 | 2 weeks |
| 30% | 6,455 | 5 | 1 week |
The 10% row is the trap. 36 days is past the four-week mark, so Benchmyrk flags the plan as low traffic. A 20% change is detectable in 2 weeks, and a 30% change in 1 week. On a store this size, test changes big enough to move the number: a new offer, a reworked product page, a different layout. A button colour won't get there.
Try your own numbers in the sample size calculator, and read the full guide for what each input means.
Stopping early without fooling yourself
A classic A/B test assumes you look once, at the end. Check it every day and stop at the first significant result, and the chance of calling a winner that isn't real climbs well above the 5% you signed up for: our simulation of 1,000 tests shows how far.
Sequential tests are designed for repeated looks. Their p-value stays valid however often you check it, so you can stop as soon as it crosses the line. The cost: they need somewhat more evidence than a classic test to call the same effect, so small effects still take the full plan.
- Visitors per version
- 10,500
- Conversion rate
- 3.0% control, 3.9% variant
- Lift and 95% interval
- 30% (13% to 50%)
- Always-valid p-value
- 0.035
- Verdict after 7 days
- Winner found
- Same rates after 5 days
- Not enough data yet
Computed with Benchmyrk's statistics engine: these inputs give these verdicts. Example numbers, not data from a store.
A 30% lift is strong enough to call after a week, a fraction of the planned sample. At five days the same rates give no verdict at all, because Benchmyrk waits for 7 full days whatever the numbers say.
- Visitors per version
- 21,000
- Conversion rate
- 3.0% control, 3.3% variant
- Classic p-value
- 0.078
- Always-valid p-value
- 0.628
- Verdict after 14 days
- Still collecting evidence
Computed with Benchmyrk's statistics engine. Example numbers, not data from a store.
A 10% lift may well be real, but after two weeks there isn't enough evidence to say so. The test keeps going toward its plan, or you decide the change isn't worth 36 days.
When to stop without a winner
- The plan is reached and nothing is significant. That's an answer: the change, if it does anything, probably does less than the effect you planned for. Keep the simpler version.
- The test passes four weeks with no end in sight. Stop, keep the control, and test something bigger or on a busier page.
- Something broke. A guardrail tripped, the split is off (a sample ratio mismatch), or a version stopped rendering. Fix it and restart; don't patch a running test and keep its data.
- The store changed under it. A sitewide sale, a new theme, a price change or a large ad campaign that started mid-test. The result may not hold once things are back to normal.
Shopify things that stretch or spoil a test
- Promotions. Shoppers who arrive for a sale behave differently. Avoid starting a test the week before a big promotion, and don't let one run across Black Friday unless the sale is what you're testing.
- Theme publishes and app installs. A theme update or a new app can change the pages being tested. Schedule them between tests.
- Email and ad bursts. A campaign day can bring a different crowd. Over whole weeks this evens out; over three days it doesn't.
- Revenue goals. Revenue per visitor varies far more than conversion, so it needs more visitors than the conversion plan says. Plan for longer when revenue or profit decides the test.
- Refunds. Orders can come back weeks later. For offer and threshold tests, compare refunds per version before rolling out.
Cheat sheet
| Situation | What to do |
|---|---|
| Fewer than 7 full days | Keep going, whatever the numbers say. |
| A sequential test calls a winner after a week or more | You can stop. Check the interval: a wide one means the size of the lift is still uncertain. |
| A classic test looks significant before the plan | Keep going. Stopping now inflates false winners. |
| The plan is reached, nothing is significant | Stop and keep the simpler version. |
| Past four weeks, no result | Stop, keep the control, and test a bigger change. |
| The split is off, or a guardrail tripped | Investigate before you read anything else. |
Sources
- Johari, Koomen, Pekelis and Walsh, "Always Valid Inference: Continuous Monitoring of A/B Tests", Operations Research, 2022
- Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020
- Fleiss, Levin and Paik, Statistical Methods for Rates and Proportions, 3rd edition, Wiley, 2003