A/B Test Sample Size Calculator
Estimate how many participants each variant needs before you start a conversion experiment. Free, no account, and nothing you enter leaves your browser.
New to the terms below? The free lesson explains them with a worked example. Learn how power and sample size work
Result
Updates as you type
- Participants per variant
- 31,234
- 31,234 in variant A and 31,234 in variant B
- Participants in total
- 62,468
- A and B together
What this estimate assumes
- Conversion
- 5% → 5.5%
- Effect
- +10% relative, +0.5 percentage points
- Significance level
- 5%, two-sided
- Power
- 80%
- Allocation
- 50/50
- Method
- Normal approximation for two independent proportions, fixed sample size, no continuity correction
How to read this
- If the real effect is at least as large as you assumed and the assumptions hold, a test of this size detects it with 80% probability. This is not the probability that B wins.
- Fix the sample size before you start and analyse once it is reached. Do not stop the experiment just because an interim p-value became significant.
- This is an estimate from an approximation, not an exact number or a guarantee.
How the calculator works
The calculator plans a classic A/B test of a conversion rate: two variants, the same number of participants in each, one yes-or-no outcome per participant, and one analysis at the end.
It uses the normal approximation for two independent proportions with a two-sided test and no continuity correction. Four things drive the result: the baseline conversion, the smallest effect you want to detect, the significance level and the power.
The effect matters most. Halving the effect you want to detect roughly quadruples the sample, which is why a realistic minimum effect is the most important input.
Worked example
A checkout converts 5% of visitors. The team wants to detect a 10% relative uplift, which means 5.5%, a difference of 0.5 percentage points. With a 5% significance level and 80% power, each variant needs 31,234 participants, 62,468 in total.
With 1,000 eligible participants per day across both variants, recruiting takes about 63 days. If that is too long, the honest options are a bigger change that could produce a bigger effect, a page with more traffic, or a metric with a higher baseline.
Note the units: a 10% relative uplift from 5% is 5.5%. A 10 percentage-point increase from 5% is 15%, a very different experiment.
Assumptions and limits
- One binary outcome per participant: converted or not. Revenue and other continuous metrics need a different method.
- Participants are randomized independently and counted once. Page views, sessions and repeat visits are not participants.
- Fixed sample size with one analysis at the end. Sequential testing and early stopping need a different design.
- Two variants with a 50/50 split. More variants or unequal splits are not covered.
- The result comes from an approximation. With very small or very large conversion rates and few participants it does not hold, and the calculator shows no number.
- The time estimate covers recruitment only. It does not include the delay between entering the test and converting.
Questions and answers
- Is the result per variant or in total?
- Both are shown. The main number is per variant: each of A and B needs that many participants. The total is twice that number.
- What is the difference between relative uplift and percentage points?
- Relative uplift is a share of the baseline: +10% from 5% gives 5.5%. Percentage points are added to the baseline: +10 points from 5% gives 15%. Mixing them up changes the sample size many times over, so the calculator always shows the change in both units.
- What does 80% power mean?
- If the real effect equals the one you entered, about 80 of 100 such experiments would show a statistically significant result. The other 20 would miss it. Power says nothing about the chance that the new variant is better.
- Which daily traffic number should I enter?
- New participants per day who can enter this experiment, for both variants together. Exclude people who are not eligible or not included in the test, and count each person once.
- Why do other calculators show a different number?
- Calculators differ in the formula, one-sided or two-sided tests, continuity corrections and rounding. Compare results only when the assumptions match. A difference of a few percent between methods is normal.
Related tools
- A/B/n test results calculator
Reading the results of a finished test with three to six variants.
- Minimum detectable effect calculator
Finding the smallest effect a test can detect with the traffic and time you have.