
A/B Testing Service-Level Agreement (Per-User) Calculator
Estimate the per-variation user sample, total traffic, and test duration needed for an A/B test to detect a target conversion lift.
Overview
Use this A/B testing service-level agreement calculator to estimate how many qualifying users each variation needs before you can reliably assess a target conversion lift. Enter your baseline conversion rate, the smallest relative lift that matters, confidence and power settings, and expected daily test traffic.
How it works
The calculator estimates sample size for a two-variation A/B test that compares conversion rates. It converts the target relative lift into an expected variant conversion rate, then uses the baseline rate, target rate, selected confidence threshold, and statistical power to estimate the users needed in each group. The total sample is divided by the expected daily test traffic to estimate duration. Smaller detectable lifts, lower baseline conversion rates, higher confidence, and higher power generally require more users.
How to use this calculator
- 1Enter the current conversion rate for the control experience.
- 2Set the minimum relative conversion lift that would be meaningful to detect.
- 3Choose a confidence level and statistical power for the experiment.
- 4Enter the daily number of eligible users and the share of traffic available to the test.
- 5Review the required users per variation and estimated test duration when setting the test SLA.
Example Calculation
Baseline conversion rate
10%
Minimum detectable lift
10%
Significance level
1.96
Statistical power
0.842
Daily eligible users
2000
Traffic sent to test
100%
Required users per variation
14,756 users
With a 10% baseline conversion rate, a target lift to 11%, 95% confidence, and 80% power, the test needs about 14,750 users per variation (roughly 29,500 total), which is about 15 days at 2,000 eligible users per day.
Frequently asked questions
What is a per-user A/B testing SLA?
It is a planning target for the number of qualifying users each variation should receive before the experiment is evaluated. It helps teams align expectations on traffic and test duration.
Why does a smaller target lift require more users?
Small differences are harder to distinguish from normal random variation. Detecting a smaller lift therefore needs a larger sample.
What confidence level should I choose for an A/B test?
A two-sided 95% confidence level is a common planning choice. Higher confidence reduces the chance of a false positive but increases the required sample.
What does statistical power mean?
Power is the chance that the test will detect the target effect when that effect truly exists. Higher power reduces the chance of missing a meaningful result but needs more users.
Does this calculator work for revenue or average order value?
This version is designed for binary conversion metrics, such as sign-ups or purchases. Revenue and other continuous metrics need a calculation that also considers metric variability.
Can I stop the test as soon as the sample target is reached?
The sample target is one important condition, but you should also check data quality, tracking consistency, and whether the experiment covered a representative time period.
Explore Related Calculators
Assumptions and warnings
Assumptions
- The experiment has two equally sized variations: one control and one variant.
- The selected significance level is two-sided, meaning the variant could perform either better or worse than the control.
- Eligible-user traffic and conversion behavior are assumed to remain reasonably stable during the test.
- The calculation uses a standard approximation for comparing two independent conversion rates.
- Results are planning estimates and do not account for data-quality issues, repeated peeking, or multiple simultaneous comparisons.
Warnings
- This calculator provides an experiment-planning estimate, not a guarantee of a statistically valid business decision.
- Do not make high-impact decisions from incomplete, biased, or improperly tracked experiment data.