What A/B test significance is and when to use it
A/B test significance testing answers one question: is the difference in conversion rate between your two variants (A and B) likely a real effect, or could it just be random noise from having finite sample sizes? This calculator runs a two-proportion z-test, the standard statistical method for comparing two conversion rates, and reports a z-score and p-value rather than an arbitrary rule of thumb. A previous version of this page flagged "significance" using only the size of the lift and a fixed visitor-count cutoff — that approach could call a result significant when it was actually noise, or dismiss a real effect that happened to occur on smaller traffic. This version instead calculates the actual probability of seeing your observed difference by chance.
Use it after running a controlled experiment where visitors were randomly split between a control (A) and a variant (B), and you tracked both total visitors and conversions for each. It applies to landing page tests, pricing experiments, email subject line tests, or any binary outcome (converted / did not convert) split test.
The formula
The test compares the two observed conversion rates using a pooled standard error:
- p̂A = Conversions A ÷ Visitors A and p̂B = Conversions B ÷ Visitors B — the observed conversion rates.
- Pooled rate p̂ = (Conversions A + Conversions B) ÷ (Visitors A + Visitors B) — the combined conversion rate assuming no real difference between groups.
- Standard error SE = √[p̂(1−p̂)(1/Visitors A + 1/Visitors B)]
- Z-score = (p̂B − p̂A) ÷ SE, converted to a two-tailed p-value using the standard normal distribution.
A result is typically called statistically significant when the p-value is below 0.05 (equivalent to a z-score beyond ±1.96), meaning there is less than a 5% chance the observed difference occurred by random chance alone if there were truly no difference between A and B.
Worked example
Suppose Variant A had 1,000 visitors and 100 conversions (10.00%), and Variant B had 1,000 visitors and 130 conversions (13.00%). The pooled rate is (100+130)/(1000+1000) = 230/2000 = 0.115. The standard error is √[0.115 × 0.885 × (1/1000 + 1/1000)] ≈ 0.01427. The z-score is (0.13 − 0.10) / 0.01427 ≈ 2.10, which corresponds to a two-tailed p-value of approximately 0.0355. Since 0.0355 is below 0.05, the calculator reports this result as statistically significant at 95% confidence — matching what it displays for these exact inputs.
Common mistakes and how to interpret the result
- Peeking early and stopping as soon as it looks significant. Checking results repeatedly and stopping the moment p dips below 0.05 inflates your false-positive rate — decide your sample size or test duration in advance and stick to it.
- Confusing a large relative lift with statistical significance. A 30% relative lift on very small samples can easily be noise (a high p-value), while a modest 5% lift on a huge sample can be highly significant — always look at the p-value and z-score, not just the lift percentage.
- Treating p ≥ 0.05 as proof of "no difference." A non-significant result often means you don't have enough data yet, not that the variants perform identically — consider running the test longer before concluding there's no effect.
- Running many simultaneous comparisons without adjustment. Testing many metrics or variants at once increases the chance of a false positive by chance alone; be more cautious interpreting significance when you are running multiple comparisons.