Statistical Significance tells you whether your results are real or just random noise. A statistically significant result means the difference between control and variant is unlikely to be due to chance alone.
What Does Statistical Significance Mean?
When a test reaches statistical significance, you can be confident the observed difference is real. The standard threshold is 95% confidence, meaning there's only a 5% probability the result occurred by chance.
95% significance means:
- 95% confident the difference is real
- 5% chance it's random variation
- Industry standard threshold
99% significance means:
- Higher confidence, stricter threshold
- Used for high-stakes decisions
- Requires more data to achieve
How It's Calculated
Statistical significance depends on:
Sample size: More visitors means more reliable results. Small samples produce noisy data.
: Lower base conversion rates need larger samples to detect changes.
Effect size: Larger differences between variants are easier to detect than small ones.
Variance: Consistent behavior is easier to measure than erratic patterns.
Reaching Significance
Tests reach significance when enough data accumulates to rule out chance. This typically requires:
- Hundreds to thousands of conversions per variant
- Days to weeks of test runtime
- Stable traffic patterns
Common Mistakes
-
Stopping tests early: Results fluctuate before stabilizing. Early winners often don't hold.
-
Peeking too often: Checking results repeatedly increases false positive risk. Set a review schedule.
-
Ignoring practical significance: A statistically significant 0.1% lift may not matter for business. Consider effect size.
-
Confusing confidence with certainty: 95% confidence still means 1 in 20 tests will show false positives.
Statistical vs Practical Significance
Statistical significance: The result is not due to chance.
Practical significance: The result matters for your business.
A test can be statistically significant but practically meaningless if the lift is too small to impact revenue.
