A/B testing terms, explained plainly.
53 terms ecommerce teams run into when they test, each with a short definition, an example and how Benchmyrk handles it.
A
- A/A test
- A test where both arms show the same page, used to check that the split, tracking and statistics behave before trusting real tests.
- A/B test
- An experiment that shows two versions of a page to randomly split visitors and compares which one performs better on a chosen goal.
- A/B/n test
- An A/B test with more than one challenger: the control is compared with two or more variants in the same experiment.
- Attribution window
- How long after a visitor first sees a test their conversions still count towards it.
- Audience
- A group of visitors defined by rules, such as device, country, traffic source or behaviour, that a test or personalization targets.
- Average order value (AOV)
- Revenue divided by the number of orders: how much a typical order is worth.
B
- Bayesian A/B testing
- Analysing a test by updating a probability model with the data, so results read as probabilities such as the chance a variant beats the control.
- Bucketing
- How a testing tool assigns each visitor to a variant and keeps them there on later visits.
C
- Chance to beat control
- The probability, given the data so far, that a variant's true rate is higher than the control's.
- Confidence interval
- A range of values for the true effect that is consistent with the data, at a stated confidence level such as 95%.
- Control
- The unchanged version in a test, which every variant is compared against.
- Conversion rate
- The share of visitors who completed a goal, such as placing an order, out of all visitors in a variant.
- Credible interval
- The Bayesian counterpart of a confidence interval: a range that contains the true value with a stated probability, given the model and the data.
- CUPED
- A variance-reduction method that uses each visitor's behaviour before a test to make results more precise.
E
- Expected loss
- In Bayesian testing, how much you would lose on average by choosing a variant if it turned out not to be the best.
- Exploratory analysis
- Looking at test results in ways that weren't planned, such as by segment, to find ideas rather than to confirm a result.
F
- False negative
- Missing a real effect: the test ends without a winner although the variant really was better.
- False positive
- Calling a winner that isn't real: the test finds a difference although the variant made none.
- Feature flag
- A switch in your code, controlled from outside, that turns a feature on or off or serves a value to a chosen share of users.
- Flicker
- When a visitor briefly sees the original page before a test's change replaces it.
- Frequentist A/B testing
- Analysing a test with p-values and confidence intervals, which describe how surprising the data would be if the variant made no difference.
G
- Guardrail metric
- A metric a test must not make worse, watched alongside the goal the test is trying to improve.
H
- Holdout
- A share of visitors kept on the original experience, so the effect of a personalization or a programme of changes can be measured.
- Holm correction
- A way to adjust p-values when a test compares several variants with the control, so the chance of any false winner stays at the chosen level.
L
- Lift
- How much better or worse a variant performs than the control, as a percentage of the control's value.
M
- Minimum detectable effect (MDE)
- The smallest lift a test is planned to detect reliably; it sets how many visitors the test needs.
- Multi-armed bandit
- A test that shifts traffic toward the better-performing variant while it runs, instead of keeping a fixed split.
- Multi-page test
- One test that changes several pages together and keeps each visitor in the same version across all of them.
- Multiple comparisons
- The problem that the more comparisons you make in one test, the more likely one looks significant by chance.
- Multivariate test (MVT)
- A test that changes several page elements at once and tries every combination, to see which version of each element works best.
- Mutually exclusive tests
- Tests set up so that no visitor is in more than one of them, to keep tests on the same pages from affecting each other.
N
- Novelty effect
- A temporary change in behaviour because something is new, which fades once visitors get used to it.
O
- Offer test
- A test that gives different groups of shoppers different offers, such as 10% off versus $10 off, and compares what each earns after the discount.
P
- Peeking
- Checking a test's results while it runs and stopping as soon as they look significant, which inflates false winners with classic p-values.
- Personalization
- Showing different visitors different versions of a page based on who they are or what they do, instead of the same page for everyone.
- Price test
- A test where different groups of shoppers see and pay different prices for the same product, to find the price that earns the most.
- Primary goal
- The one metric a test is judged on, chosen before it starts.
- Profit per visitor
- Gross profit from orders divided by visitors: revenue minus product costs and other variable costs, per visitor in a variant.
- P-value
- The probability of seeing a difference at least as large as the observed one if the variant actually made no difference.
R
- Revenue per visitor (RPV)
- Total order revenue divided by visitors in a variant: a goal that combines how many visitors buy and how much they spend.
S
- Sample ratio mismatch (SRM)
- When the number of visitors in each variant differs from the planned split by more than chance can explain: a sign the test is broken.
- Sample size
- How many visitors a test needs per variant to detect the minimum detectable effect with the chosen confidence and power.
- Segment
- A subset of a test's visitors, such as mobile visitors or new visitors, whose results can be looked at separately.
- Sequential testing
- Statistical methods that let you check a test's results as often as you like and stop early, without inflating false winners.
- Server-side testing
- Running experiments in your own server or edge code rather than by changing the page in the visitor's browser.
- Shipping test
- A test of shipping offers, such as the free-shipping threshold, where each group of shoppers gets its own shipping price at checkout.
- Split URL test
- A test that sends part of the traffic to a different URL, such as a redesigned landing page, and compares the outcomes.
- Statistical power
- The chance that a test detects an effect of a given size when that effect is really there.
- Statistical significance
- A result is significant when its p-value is below a chosen threshold, such as 0.05, meaning chance alone is an unlikely explanation.
T
- Template test
- A Shopify test that sends part of the traffic to an alternate page template, to compare a new layout without editing the live theme.
- Traffic allocation
- How much of the eligible traffic enters a test, and how it is divided between the variants.
V
- Variant
- A changed version of a page or experience in a test, compared against the control.
W
- Web pixel
- Shopify's sandboxed way for apps to receive storefront and checkout events, such as checkout started and purchase, without editing the theme.
Put the terms to work.
We'll show how Benchmyrk reads a result on demo data. Or start the free trial on Shopify.