
A/B Testing Statistical Significance Calculator FAQ
Answers to common questions about per-user A/B conversion tests, z-scores, evidence ratios, confidence thresholds, and limitations.
This FAQ explains what the calculator measures, which inputs it needs, and how to interpret a per-user conversion-rate comparison. It is educational information; calculator results are estimates based on the supplied data and assumptions.
General Calculator Questions
Basic information about what this calculator compares and when it fits an experiment.
What does this A/B testing calculator measure?
It compares two observed per-user binary conversion rates and reports the rate difference, relative uplift, z-score, and evidence ratio against a selected threshold.
What is a per-user conversion?
It means a unique user either completed the defined conversion during the measurement window or did not. Each user should be counted once per assigned variant.
Which metrics fit this calculator?
Binary user-level metrics such as signup completion, purchase conversion, activation, or click-through conversion can fit when users are counted uniquely.
Which metrics do not fit this calculator?
Revenue per user, average order value, session counts, repeated event totals, and other continuous or non-binary metrics require a different approach.
Inputs and Outputs
How to enter experiment data and interpret the displayed measures.
What should I enter for users?
Enter the number of unique eligible users assigned to each variant during the same analysis period.
What should I enter for conversions?
Enter the number of those unique users who completed the same defined conversion event. Conversions cannot exceed users for this per-user calculation.
What is the absolute conversion-rate difference?
It is Variant B's conversion rate minus Variant A's conversion rate. It is best read in percentage points.
What is relative uplift?
It is the absolute rate difference divided by Variant A's rate, expressed as a percentage. A negative value means Variant B's observed rate is lower.
What is a z-score in an A/B test?
It is the absolute observed rate difference divided by its estimated standard error. It indicates how large the observed gap is relative to estimated random variation.
Thresholds and Significance
How the selected confidence threshold and evidence ratio work.
What does the evidence ratio mean?
It is the z-score divided by the selected critical z-score. A value above 1.00 means the result clears the selected two-sided threshold.
Which threshold should I select?
The calculator offers 90%, 95%, and 99% two-sided thresholds. A higher confidence threshold requires a larger z-score to clear it.
Is this a one-sided or two-sided test?
It uses two-sided critical z-scores. That means it tests for a conversion-rate difference in either direction.
Does an evidence ratio below 1 mean there is no difference?
No. It means the observed data do not clear the selected threshold under this calculation. The true difference may be smaller, absent, or not precisely estimated with the available data.
Does statistical significance prove Variant B is better?
No. A two-sided significant result can indicate either an increase or a decrease. Also consider the sign and size of the observed change and the quality of the experiment.
Accuracy and Experiment Design
Important assumptions and reasons a simple calculation can differ from a real experiment decision.
Why does sample size matter?
Larger groups generally reduce the standard error, making a given conversion-rate difference easier to distinguish from ordinary random variation.
Does this calculator adjust for repeated checking?
No. Repeatedly checking results and stopping when they look favorable can increase false-positive risk.
Does this calculator account for multiple variants or metrics?
No. Testing many variants, segments, or metrics can require adjustments beyond this simple two-group calculation.
Can unequal group sizes be used?
Yes. The formula uses each variant's own user count and conversion rate, although substantially uneven allocation can reduce precision compared with balanced groups at the same total traffic.
Is a statistically significant result automatically worth implementing?
No. Statistical evidence and practical importance are different. The absolute change, likely value, costs, risks, and guardrail metrics may also matter.
What does the evidence ratio mean?
It is the observed z-score divided by the selected critical z-score. A value above 1.00 clears the selected two-sided threshold.
Explore Related Questions
Ready to see what you can calculate?
Open the calculator and get personalized results in seconds.
