Benchmyrk

Glossary

A/B testing terms, explained plainly.

53 terms ecommerce teams run into when they test, each with a short definition, an example and how Benchmyrk handles it.

A

A/A test
A test where both arms show the same page, used to check that the split, tracking and statistics behave before trusting real tests.
A/B test
An experiment that shows two versions of a page to randomly split visitors and compares which one performs better on a chosen goal.
A/B/n test
An A/B test with more than one challenger: the control is compared with two or more variants in the same experiment.
Attribution window
How long after a visitor first sees a test their conversions still count towards it.
Audience
A group of visitors defined by rules, such as device, country, traffic source or behaviour, that a test or personalization targets.
Average order value (AOV)
Revenue divided by the number of orders: how much a typical order is worth.

B

Bayesian A/B testing
Analysing a test by updating a probability model with the data, so results read as probabilities such as the chance a variant beats the control.
Bucketing
How a testing tool assigns each visitor to a variant and keeps them there on later visits.

C

Chance to beat control
The probability, given the data so far, that a variant's true rate is higher than the control's.
Confidence interval
A range of values for the true effect that is consistent with the data, at a stated confidence level such as 95%.
Control
The unchanged version in a test, which every variant is compared against.
Conversion rate
The share of visitors who completed a goal, such as placing an order, out of all visitors in a variant.
Credible interval
The Bayesian counterpart of a confidence interval: a range that contains the true value with a stated probability, given the model and the data.
CUPED
A variance-reduction method that uses each visitor's behaviour before a test to make results more precise.

E

Expected loss
In Bayesian testing, how much you would lose on average by choosing a variant if it turned out not to be the best.
Exploratory analysis
Looking at test results in ways that weren't planned, such as by segment, to find ideas rather than to confirm a result.

F

False negative
Missing a real effect: the test ends without a winner although the variant really was better.
False positive
Calling a winner that isn't real: the test finds a difference although the variant made none.
Feature flag
A switch in your code, controlled from outside, that turns a feature on or off or serves a value to a chosen share of users.
Flicker
When a visitor briefly sees the original page before a test's change replaces it.
Frequentist A/B testing
Analysing a test with p-values and confidence intervals, which describe how surprising the data would be if the variant made no difference.

G

Guardrail metric
A metric a test must not make worse, watched alongside the goal the test is trying to improve.

H

Holdout
A share of visitors kept on the original experience, so the effect of a personalization or a programme of changes can be measured.
Holm correction
A way to adjust p-values when a test compares several variants with the control, so the chance of any false winner stays at the chosen level.

L

Lift
How much better or worse a variant performs than the control, as a percentage of the control's value.

M

Minimum detectable effect (MDE)
The smallest lift a test is planned to detect reliably; it sets how many visitors the test needs.
Multi-armed bandit
A test that shifts traffic toward the better-performing variant while it runs, instead of keeping a fixed split.
Multi-page test
One test that changes several pages together and keeps each visitor in the same version across all of them.
Multiple comparisons
The problem that the more comparisons you make in one test, the more likely one looks significant by chance.
Multivariate test (MVT)
A test that changes several page elements at once and tries every combination, to see which version of each element works best.
Mutually exclusive tests
Tests set up so that no visitor is in more than one of them, to keep tests on the same pages from affecting each other.

N

Novelty effect
A temporary change in behaviour because something is new, which fades once visitors get used to it.

O

Offer test
A test that gives different groups of shoppers different offers, such as 10% off versus $10 off, and compares what each earns after the discount.

P

Peeking
Checking a test's results while it runs and stopping as soon as they look significant, which inflates false winners with classic p-values.
Personalization
Showing different visitors different versions of a page based on who they are or what they do, instead of the same page for everyone.
Price test
A test where different groups of shoppers see and pay different prices for the same product, to find the price that earns the most.
Primary goal
The one metric a test is judged on, chosen before it starts.
Profit per visitor
Gross profit from orders divided by visitors: revenue minus product costs and other variable costs, per visitor in a variant.
P-value
The probability of seeing a difference at least as large as the observed one if the variant actually made no difference.

R

Revenue per visitor (RPV)
Total order revenue divided by visitors in a variant: a goal that combines how many visitors buy and how much they spend.

S

Sample ratio mismatch (SRM)
When the number of visitors in each variant differs from the planned split by more than chance can explain: a sign the test is broken.
Sample size
How many visitors a test needs per variant to detect the minimum detectable effect with the chosen confidence and power.
Segment
A subset of a test's visitors, such as mobile visitors or new visitors, whose results can be looked at separately.
Sequential testing
Statistical methods that let you check a test's results as often as you like and stop early, without inflating false winners.
Server-side testing
Running experiments in your own server or edge code rather than by changing the page in the visitor's browser.
Shipping test
A test of shipping offers, such as the free-shipping threshold, where each group of shoppers gets its own shipping price at checkout.
Split URL test
A test that sends part of the traffic to a different URL, such as a redesigned landing page, and compares the outcomes.
Statistical power
The chance that a test detects an effect of a given size when that effect is really there.
Statistical significance
A result is significant when its p-value is below a chosen threshold, such as 0.05, meaning chance alone is an unlikely explanation.

T

Template test
A Shopify test that sends part of the traffic to an alternate page template, to compare a new layout without editing the live theme.
Traffic allocation
How much of the eligible traffic enters a test, and how it is divided between the variants.

V

Variant
A changed version of a page or experience in a test, compared against the control.

W

Web pixel
Shopify's sandboxed way for apps to receive storefront and checkout events, such as checkout started and purchase, without editing the theme.

Put the terms to work.

We'll show how Benchmyrk reads a result on demo data. Or start the free trial on Shopify.