
A/B Testing Service-Level Agreement Calculator FAQ
Answers to common questions about A/B test traffic requirements, sample size, statistical settings, and estimated duration.
This FAQ explains how to use a per-user A/B testing SLA estimate for a two-variation conversion-rate experiment. Results are planning estimates and should be considered alongside sound experiment design and data-quality checks.
General A/B testing SLA questions
Understand what the calculator is designed to estimate.
What is an A/B testing service-level agreement?
In this context, it is a planning target for the qualifying users and approximate time needed before evaluating a conversion experiment.
What does per-user mean in this calculator?
The key output is the number of qualifying users required in each variation, rather than only a total traffic figure.
What metric is this calculator for?
It is for binary conversion metrics, where each eligible user either converts or does not convert.
Does the calculator support more than one variant?
No. It estimates a single control-versus-variant comparison with equal group sizes.
Inputs and formula
Learn what each input changes in the calculation.
What is baseline conversion rate?
It is the expected control-group conversion rate before the test, entered as a percentage.
What is minimum detectable lift?
It is the smallest relative improvement over the baseline that the test is designed to detect.
What is the difference between relative lift and percentage points?
A 10% relative lift on a 10% baseline gives an 11% variant rate, which is a one-percentage-point absolute increase.
What does statistical power mean?
Power is the probability of detecting the specified target effect if it is genuinely present under the calculation assumptions.
Why is the significance setting two-sided?
A two-sided setting tests for a meaningful difference in either direction, whether the variant is higher or lower than the control.
Traffic and duration
See how traffic availability is used to estimate runtime.
How is estimated test duration calculated?
Total required users are divided by daily eligible users multiplied by the share of traffic sent to the test, then rounded up.
Why is eligible traffic different from all traffic?
Only users who can enter the experiment and contribute a valid metric observation count toward the expected sample.
Will sending less traffic to the test change users per variation?
It does not change the statistical sample estimate, but it reduces daily test traffic and lengthens the estimated duration.
Should I use average daily traffic?
Use a realistic average of qualifying traffic, while recognizing that traffic can vary over time.
Accuracy and interpretation
Know what the result does and does not cover.
Is the sample-size result guaranteed?
No. It is an approximation based on the entered assumptions and cannot guarantee a particular business or statistical outcome.
Can I use this for revenue per user?
No. Revenue is usually a continuous or highly skewed metric and needs a method that accounts for its variability.
Does reaching the sample size make the result valid automatically?
No. The experiment also needs reliable assignment, accurate tracking, and an analysis approach consistent with the test plan.
Can I repeatedly check results during the test?
Repeated monitoring can affect interpretation unless the experiment design and analysis method account for interim looks.
Why does a smaller target lift require more A/B test users?
Smaller conversion differences are harder to separate from normal random variation, so they require more data.
Explore Related Questions
Ready to see what you can calculate?
Open the calculator and get personalized results in seconds.
