P-value
The probability of seeing a difference at least as large as the observed one if the variant actually made no difference.
A small p-value means the data would be surprising if there were no effect. Below the chosen level (usually 0.05), the result is called significant.
It is not the probability that the variant is better, and it doesn't measure the size of the effect. A classic p-value also assumes one look at a planned sample; checking it every day and stopping early inflates false winners.
A test ends with p = 0.012 for the new product photos: a difference this large would happen about 1% of the time if the photos made no difference.
Related terms
- Statistical significanceA result is significant when its p-value is below a chosen threshold, such as 0.05, meaning chance alone is an unlikely explanation.
- Sequential testingStatistical methods that let you check a test's results as often as you like and stop early, without inflating false winners.
- PeekingChecking a test's results while it runs and stopping as soon as they look significant, which inflates false winners with classic p-values.
- Confidence intervalA range of values for the true effect that is consistent with the data, at a stated confidence level such as 95%.