A/B Test Results Calculator
Enter participants and conversions for two variants of a finished test. You get the difference in percentage points, the relative lift, a confidence interval and a p-value, with a plain reading of what they do and do not say.
New to reading test results? The free lesson walks through it step by step. Learn how conversion rate tests work
Testing more than two variants? Use the A/B/n results calculator
Result
Enter the counts of both variants and press Analyze results. The example is ready to try.
How the calculator works
You enter how many people saw each variant and how many of them converted. The calculator turns the counts into two conversion rates and their difference, B minus A.
A two-sided test asks whether a difference this large is likely if both variants convert equally. If the p-value is below the significance level you chose, the result is called statistically significant.
The confidence interval shows how large or small the true difference could plausibly be. A narrow interval far from zero is strong evidence; a wide interval that crosses zero means the test cannot tell the variants apart yet.
Want the p-value explained with a picture, or a one-sided test? Use the statistical significance calculator.
Worked example
A test ran with 10,000 participants in each variant. A (control) had 500 conversions, B had 600. The rates are 5% and 6%.
The difference is +1 percentage point, which is a relative lift of +20% over A. The 95% interval runs from about +0.37 to +1.63 percentage points, and the p-value is about 0.0019.
So B converted better in this test, and a difference this large would be unusual if the variants were equal. Whether +1 point is worth shipping is a business question: weigh it against costs and guardrail metrics.
Method and formulas
Rates are pA = xA / nA and pB = xB / nB, where x is conversions and n is participants. The difference is d = pB − pA; the relative lift is d / pA.
The test pools both variants to estimate the rate under "no difference", then compares d with its standard error. The p-value comes from the normal distribution, two-sided.
The interval for d is Newcombe's method: it combines a Wilson score interval for each variant. It stays sensible near 0% and 100%, where the simple "estimate ± 1.96 × standard error" does not.
pA = xA / nA, pB = xB / nB, d = pB − pA
pPool = (xA + xB) / (nA + nB)
SE0 = √(pPool (1 − pPool) (1/nA + 1/nB)), z = d / SE0
p = erfc(|z| / √2)
lower = d − √((pB − L_B)² + (U_A − pA)²)
upper = d + √((U_B − pB)² + (pA − L_A)²)Both methods are approximations. They need at least 10 conversions and 10 non-conversions in each variant; below that the calculator shows only the observed rates.
Questions and answers
- Should I count people, sessions or events?
- People. Each participant counts once, and a conversion means that person converted at least once. Counting sessions or repeated purchases breaks the assumption that observations are independent and makes results look more certain than they are.
- What is the difference between percentage points and relative lift?
- Percentage points subtract one rate from the other: 6% − 5% = 1 point. Relative lift divides that by the control: 1 / 5 = +20%. Both describe the same result, so always say which one you mean.
- The result is not significant. Does that mean the variants are the same?
- No. It means this test did not detect a difference. A real difference may be smaller than the test could see. Look at the interval: if it still includes differences you care about, the test was too small to decide.
- Can I check the results every day and stop when they turn significant?
- Not with this method. It assumes one analysis at a planned point. Checking repeatedly and stopping at the first significant result makes false wins much more likely than the significance level suggests.
- Why is the relative lift "not defined"?
- Relative lift divides by the control rate. When the control has no conversions, that division is impossible, so the calculator shows the difference in percentage points only instead of an infinite percentage.
- Why is there no p-value for my small test?
- With fewer than 10 conversions or non-conversions in a variant, the approximation behind the test and the interval becomes unreliable. The calculator shows the observed rates and says so, rather than a number that looks precise and is not.
Terms used here
Related tools
- A/B/n test results calculator
Reading the results of a finished test with three to six variants.
- A/B test sample size calculator
Planning for two variants (A/B) before the test starts.
- Confidence interval calculator
Seeing how precise a single conversion rate is.