Test types

Send more traffic to what's winning, while the test runs.

For promotions and short campaigns: Benchmyrk moves traffic toward the better variant and keeps a minimum share on every one.

Fixed split or auto-allocate?

  • Best for

    Fixed split

    A clear answer about what works, to keep and build on.

    Auto-allocate

    Promotions and short campaigns, where revenue during the test matters most.

  • Traffic

    Fixed split

    The split you set, for the whole test.

    Auto-allocate

    Even at first, then moved toward the variant most likely to be best.

  • Results

    Fixed split

    Chance to beat control, 95% intervals and p-values.

    Auto-allocate

    Each variant's chance of being best and expected loss; no p-values.

  • Speed to a firm answer

    Fixed split

    Faster: every variant keeps its full share.

    Auto-allocate

    Slower: weaker variants get less traffic, so they take longer to rule out.

How traffic is splitDemo data

How the split moves

  1. 01

    Warm-up

    An even split until the warm-up has passed (24 hours by default, 6 to 72) and every variant has 100 visitors, plus 5 buyers on revenue goals.

  2. 02

    Hourly updates

    Each hour the split moves toward each variant's chance of being best, by at most 10 points per variant.

  3. 03

    A minimum share

    Every variant keeps at least 10% of visitors by default (5% to 25%), so the evidence keeps coming in.

  4. 04

    Returning visitors stay put

    A visitor keeps the variant they saw first; only new visitors follow the new split.

A leader is called only when it's clear

Checking "chance of being best" every hour finds false leaders: in our simulated A/A tests, a 95% rule picked a leader in about half of the runs where the variants were equal. So Benchmyrk calls a leader only when its anytime-valid confidence sequence lies above every other variant's, which holds however often anyone looks.

  • In the same simulation the rule called no false leader in 300 of 300 runs.
  • The sample ratio check accounts for the changing split, so a moving split isn't mistaken for a broken one.
  • Guardrails, staged rollouts, mutually exclusive groups and approvals work as for any test.

Good to know

  • Works with A/B, split URL and multi-page tests.
  • CUPED variance reduction isn't available with auto-allocation.
  • The model assumes each variant's rate is stable during the test. Strong time-of-day or weekday swings can bias results; the warm-up and the 10-point cap limit this.
  • Profit goals hold the split until product costs are complete enough to judge profit.

Questions

Why are there no p-values for auto-allocated tests?

The split depends on the results so far, which biases p-values and frequentist intervals. The Bayesian estimates stay valid, so results show each variant's chance of being best, its credible interval and expected loss.

When should I use a fixed split instead?

When a firm answer matters more than revenue during the test. Auto-allocation spends most traffic on the leader, so it takes longer to rule out the weaker variants.

Can I change the minimum share or warm-up?

Yes, before the test starts: a minimum share of 5% to 25% per variant and a warm-up of 6 to 72 hours.

Running a promotion soon?

We'll show how auto-allocation behaves on demo data and when a fixed split is the better choice. Or install the Shopify app and try it.