A/B/n Test Results Calculator
Compare three to six variants of a finished conversion test. The calculator checks every pair of variants and corrects for the number of comparisons, so an extra variant does not quietly raise the risk of a false win.
Why is comparing many variants risky? The free lesson explains it with examples. Understand multiple testing mistakes
Result
Enter the counts of each variant and compare them. The example above is ready to try.
How the calculator works
Enter participants and conversions for each variant. The calculator compares every pair of variants: with three variants that is three comparisons, with four it is six, with six it is fifteen.
Each pair gets a two-sided test for a difference in conversion rates and an interval for that difference. Because several comparisons are made at once, the chance of at least one false finding grows. The Bonferroni correction compensates: p-values are multiplied by the number of comparisons, and intervals are widened to match.
The set of comparisons is fixed by the number of variants. You cannot drop a pair after seeing the results, because choosing comparisons afterwards brings the false findings back.
Worked example
Three variants, 10,000 participants each. A (control) has 500 conversions, B has 600 and C has 620. The observed rates are 5%, 6% and 6.2%, so C has the highest observed conversion.
After the correction, B differs from A (adjusted p = 0.0058) and C differs from A (adjusted p = 0.0007). But C does not differ from B: adjusted p = 1, and the interval for the difference runs from −0.61 to 1.01 percentage points.
So both alternatives beat the control, and the data cannot tell B and C apart. That does not make them equal or joint winners. It means this test gives no statistical reason to prefer C over B.
Assumptions and limits
- One binary outcome per participant: converted or not. Revenue and other continuous metrics need a different method.
- Participants are randomized independently, belong to one variant only and come from the same test and the same time window.
- The test is finished and analysed once. Checking repeatedly until something is significant is not covered by the correction.
- All pairs of three to six variants. Comparing with the control only, or other corrections, are not offered.
- The calculator does not check whether the split between variants worked. Equal or unequal group sizes prove nothing about that.
- With fewer than 10 conversions or 10 non-conversions in any variant, no statistical comparison is shown.
Questions and answers
- The highest rate is clear. Why is there no leader?
- The highest observed rate is a fact about this sample. To be the leader, a variant also needs a statistically significant difference from every other variant after the correction. Two close variants often cannot be told apart even when both beat the control.
- I only care about the control. Why do all pairs count?
- Naming the best variant means comparing the alternatives with each other as well. The calculator fixes the full set of pairs in advance, so the correction matches the question. A control-only comparison is a different analysis and is not offered here.
- What does the correction not fix?
- It accounts for the comparisons in this one test and metric. It does not account for looking at the results many times, for testing several metrics or segments, or for variants that were run and left out.
- Why are some results not calculated?
- The method relies on an approximation that needs at least 10 conversions and 10 non-conversions in every variant. Below that the numbers would look precise and be unreliable, so the calculator shows only the observed rates.
- Why not run separate A/B calculations for each pair?
- Each separate test at 5% has its own 5% risk of a false finding. Across three pairs the risk of at least one is already about 12%, and with six variants it is several times the 5% you intended. Separate calculators do not know about each other, so they cannot correct for that.
Terms used here
Related tools
- A/B test sample size calculator
Planning for two variants (A/B) before the test starts.
- Minimum detectable effect calculator
Finding the smallest effect a test can detect with the traffic and time you have.