A/B/n testing: how to compare three or more variants without a false winner
A/B/n testing multiplies the chances of a false winner. How the Bonferroni and Holm corrections work, with a four-variant example and the traffic cost.
By Sergey BruhPublished 9 min read
In A/B/n testing you compare a control with two or more alternatives at once, and every extra variant adds comparisons, each of which is one more chance for a false positive, so the variant with the lowest p-value cannot simply be declared the winner. With four variants there are six possible pairs. Test each pair at the usual 5% level and the chance of at least one false positive is about 20–26% instead of 5%. The fix is to decide in advance which comparisons you will make and to apply a correction (Bonferroni or Holm) that keeps the risk for the whole test at 5%. It works, and it costs traffic.
Four buttons and one tempting winner
CatChow, an online pet-food shop, tests the label on its checkout button. Emily, the product manager, had three new ideas, so all three went into the test next to the current label. It ran for four weeks at about 3,000 checkout visitors a day, split equally:
Emily runs a two-proportion z-test for every pair and gets four "significant" results out of six. D has the highest conversion and the lowest p-value against the control, so the slide almost writes itself: "D wins, +19.8%". Two questions come first. How many of those four results could be noise? And is D really better than C?
Why the risk grows with every comparison
One comparison at has a 5% chance of a false positive when there is no real difference. With independent comparisons, the chance that at least one of them fires is:
The formula assumes independent comparisons. Inside one test they share data (every pair with A uses the same control group), so the real risk is somewhat lower: a simulation gives about 12% for all pairs of three variants and about 20% for all pairs of four. Either way it is several times the 5% you agreed to. This is the multiple comparisons problem: the lowest of six p-values means much less than the same p-value in a plain A/B test.
The two-step approach
Step 1: an overall test. Is there any difference between the variants at all? For conversion rates this is a chi-square test on the table of orders and non-orders for all four variants. CatChow's data gives with 3 degrees of freedom and . The variants are not all the same, but the test does not say which ones differ. For averages such as order value, the same logic is ANOVA followed by post hoc tests.
Step 2: pairwise comparisons with a correction. Compare the pairs, but hold each p-value to a stricter bar, so that the risk of at least one false positive across the whole family of comparisons stays at 5%. The protection comes from this step: the correction works with or without an overall test.
Bonferroni and Holm in plain language
Bonferroni splits the 5% evenly. With six comparisons each one has to pass . Equivalently, multiply every p-value by 6 (capping the result at 1) and compare it with 0.05.
Holm spends the same 5% step by step. Sort the p-values from smallest to largest, multiply the smallest by 6, the next by 5, then by 4 and so on. Walk down the list and stop at the first adjusted value of 0.05 or more: that comparison and everything after it are not significant.
Adjusted values are calculated from unrounded p-values. What the table says:
- Without a correction, four pairs pass 0.05. After Bonferroni two remain: D beats A and D beats B.
- Holm also confirms C vs A (0.0362), which Bonferroni narrowly rejects (0.0543). Both methods hold the family risk at 5%, but Holm is never stricter than Bonferroni, so it misses fewer real effects.
- D has the highest conversion, but D vs C is far from significant under either method. This test cannot tell C and D apart.
The honest summary is "D beats the current button and B; there is no evidence that D is better than C", not "D wins".
To check your own numbers, use the free A/B/n test results calculator. It takes three to six variants, compares every pair with a two-sided z-test, applies the Bonferroni correction and adds an adjusted interval for each difference. For CatChow's counts it returns the Bonferroni values from the table and the verdict that no single variant outperformed all the others. It has no overall test, Holm correction or control-only mode.
Control only or all pairs?
If the only question is "does any new label beat the current one?", three comparisons are enough: B, C and D against A. Bonferroni then multiplies by 3, and C vs A gets an adjusted p of 0.0271, which is significant.
The price is the conclusion you are allowed to draw. You may say "C and D each beat the control", but nothing about D against C: naming the single best variant needs all pairs. The family has to be fixed before launch. Choosing the smaller one after seeing the results is a quieter way of picking the lowest p-value.
What it costs in traffic
Two things get more expensive at once: the threshold is stricter, and traffic is split more ways. For a 4% baseline, a 10% relative lift (4.0% to 4.4%) and 80% power for each comparison:
Compared with a plain A/B test, four variants with all pairs need 54% more users per variant and twice as many groups: about 3.1 times the traffic, or 12 full weeks instead of 4.
So what did CatChow's four weeks buy? With 21,000 visitors per variant and a threshold of 0.0083, the test had 80% power only for lifts of about 17% or more. The same four weeks as a plain A/B test, with 42,000 visitors per group, would have caught a lift of about 10%.
The A/B test sample size calculator plans a two-variant test. For a rough A/B/n estimate, set its significance level to 1% (close to the corrected 0.83%, so slightly optimistic) and multiply the users per variant by the number of variants.
The same problem in metrics and segments
Variants are only one way to multiply comparisons. Check one A/B test on 10 metrics and the chance of at least one false positive is about 40%. Slice it by 8 segments (device, channel, new or returning customers) and it is about 34% for the segments alone. "The new button works for returning Android users" is usually the multiple comparisons problem in disguise.
Choose one primary metric in advance and treat segment findings as hypotheses for the next test. Peeking at the results every day multiplies comparisons too. A correction for variants covers none of this.
Common mistakes
Declaring the highest rate the winner. The highest observed conversion is a fact about this sample. A winner also needs a significant difference from the other variants after the correction.
Running each pair through a separate A/B calculator. Each calculation knows nothing about the others, so nothing corrects for their number.
Dropping variants after the fact. Removing B and C because they "obviously lost" and re-running D vs A as a plain A/B test is choosing comparisons after seeing the data.
Reading "not significant" as "equal". The adjusted interval for D minus C runs from −0.27 to 0.82 percentage points: D may be clearly better, or slightly worse.
What to do in practice
- Before launch, write down the variants, the primary metric, the family of comparisons (control only or all pairs) and the correction.
- Plan the sample with the corrected threshold and the real traffic split, then round up to full weeks.
- Analyse once, at the end, and report adjusted p-values with intervals.
- Prefer fewer, bolder variants. Button labels that differ by a word or two will differ from each other by a fraction of a percentage point, and telling them apart takes months. Variants that differ in substance have a chance of an effect you can detect. If traffic is limited, pick the strongest idea by judgement or customer research and run a plain A/B test.
Key takeaways
- In A/B/n testing the risk of at least one false positive grows with the number of comparisons: about 1 in 4 for six independent comparisons at 5%.
- Fix the family of comparisons and the correction before launch.
- Bonferroni multiplies every p-value by the number of comparisons. Holm does it step by step and misses fewer real effects.
- The highest conversion is not a winner until it differs significantly from the other variants.
- Four variants with all pairs need about three times the traffic of an A/B test.
FAQ
What is A/B/n testing?
A test where traffic is split between a control and two or more alternatives of the same element, for example four button labels. The "n" stands for any number of variants. A multivariate test is different: it varies several elements at once and compares their combinations.
Bonferroni or Holm: which correction should I use?
Both keep the risk for the whole family of comparisons at the chosen level. Holm is never stricter, so it is the better default for p-values. Bonferroni is simpler to explain and comes with matching adjusted intervals. Choose before you see the results.
Do I need a correction if I only compare variants with the control?
Yes. Three comparisons with the control at 5% each still give up to a 14% risk of at least one false positive. The correction is milder than for all pairs, because the family is smaller: three comparisons instead of six.
Learn it hands-on
Multiple comparisons, peeking and the novelty effect break tests that look perfectly sound. The free lesson Pitfalls: peeking, multiple testing, novelty effect walks through them on CatChow's tests, with quizzes and a practice task checked by AI. It is part of the free course Statistics for Product Managers and Marketers.