
A/B Testing Statistical Significance Calculator Examples
Worked A/B testing examples showing how conversion rates, lift, z-scores, and significance ratios change with different sample sizes.
These examples use a two-proportion z-test and a critical z-score of 1.96 unless stated otherwise. They illustrate why both the conversion difference and the number of observations affect the result.
Small test with a large-looking lift
A landing page test has 500 visitors in each group.
Input Summary
Control visitors and conversions
500 visitors, 25 conversions
Variation visitors and conversions
500 visitors, 32 conversions
Critical z-score
1.96
Calculation Breakdown
- 1Control rate25 / 500 × 1005.00%
- 2Variation rate32 / 500 × 1006.40%
- 3Absolute lift6.40% − 5.00%1.40 percentage points
- 4Z-scoreabs(0.064 − 0.050) / 0.0147 approximately0.95
- 5Significance ratio0.95 / 1.960.48×
Result Summary
Significance ratio
0.48×
A/B Testing Statistical Significance Calculator
The variation shows a 28.0% relative lift, but the approximate significance ratio is only 0.48×.
Equal-sized test that reaches a 95% threshold
An ecommerce signup flow is tested with balanced traffic.
Input Summary
Control visitors and conversions
5,000 visitors, 250 conversions
Variation visitors and conversions
5,000 visitors, 300 conversions
Critical z-score
1.96
Calculation Breakdown
- 1Control rate250 / 5,000 × 1005.00%
- 2Variation rate300 / 5,000 × 1006.00%
- 3Lift6.00% − 5.00%; 1.00% / 5.00%1.00 percentage point; 20.0% relative
- 4Z-scoreabs(0.0600 − 0.0500) / 0.00456 approximately2.19
- 5Significance ratio2.19 / 1.961.12×
Result Summary
Significance ratio
1.12×
A/B Testing Statistical Significance Calculator
The observed 20.0% relative lift has an approximate z-score of 2.19 and reaches the 1.96 threshold.
High baseline rate with a modest improvement
A returning-user prompt is tested on 20,000 visitors per version.
Input Summary
Control visitors and conversions
20,000 visitors, 4,000 conversions
Variation visitors and conversions
20,000 visitors, 4,200 conversions
Critical z-score
1.96
Calculation Breakdown
- 1Control rate4,000 / 20,000 × 10020.00%
- 2Variation rate4,200 / 20,000 × 10021.00%
- 3Relative lift(21.00% − 20.00%) / 20.00% × 1005.0%
- 4Z-scoreabs(0.2100 − 0.2000) / 0.00405 approximately2.47
- 5Significance ratio2.47 / 1.961.26×
Result Summary
Significance ratio
1.26×
A/B Testing Statistical Significance Calculator
A 1.00 percentage point, 5.0% relative lift produces an approximate 1.26× significance ratio.
Lower conversion rate at a stricter threshold
A lead form test has 10,000 visitors in each group and uses a 99% two-sided cutoff.
Input Summary
Control visitors and conversions
10,000 visitors, 100 conversions
Variation visitors and conversions
10,000 visitors, 130 conversions
Critical z-score
2.58
Calculation Breakdown
- 1Control rate100 / 10,000 × 1001.00%
- 2Variation rate130 / 10,000 × 1001.30%
- 3Relative lift0.30% / 1.00% × 10030.0%
- 4Z-scoreabs(0.0130 − 0.0100) / 0.00148 approximately2.03
- 5Significance ratio2.03 / 2.580.79×
Result Summary
Significance ratio
0.79×
A/B Testing Statistical Significance Calculator
The variation has a 30.0% relative lift but an approximate 0.79× ratio at the 99% two-sided threshold.
How to Read Your Results
Compare the significance ratio with 1.00×: at or above 1.00× reaches the selected z-score threshold.
Use the z-score to see the strength of statistical evidence relative to the estimated sampling variation.
Read absolute lift in percentage points; it is often easier to assess operationally than relative lift.
Use relative lift to express the change compared with the control baseline, especially when baselines differ.
Review test setup, tracking quality, duration, and guardrail metrics before interpreting a threshold result as broadly useful.
Assumptions & Important Notes
- Every example treats the metric as a binary conversion outcome per eligible observation.
- Examples use pooled two-proportion z-test calculations and rounded displayed values.
- Traffic groups are assumed to be independently assigned and measured over comparable conditions.
- The calculations illustrate statistical evidence, not expected business value or a guaranteed future outcome.
Related Examples
Frequently Asked Questions
Why do two tests with the same lift have different significance ratios?
Different baseline rates and sample sizes produce different standard errors. A smaller standard error produces a higher z-score for the same observed rate difference.
Does a higher relative lift always create a higher z-score?
No. A high relative lift at a very low baseline or in a small sample can still have a low z-score.
Can a result pass at 95% but fail at 99%?
Yes. A 99% two-sided threshold uses a higher critical z-score, so it requires stronger evidence.
Why are displayed example figures approximate?
Rates, standard errors, and z-scores are rounded for readability. The calculator uses the underlying input values for its result.
What should I compare besides the significance ratio?
Consider the conversion lift, sample quality, measurement consistency, test duration, guardrail metrics, and the practical impact of the observed change.
Ready to calculate your own result?
Use the live calculator with your own inputs, timing, and preferences.