significance level (alpha)
The threshold for rejecting H0, fixed before the test, usually 0.05; equals the accepted false-positive rate.
The significance level, written α (alpha), is the line you draw before a test: if the p-value falls below it, you call the result statistically significant. The usual choice is α = 0.05, or 5%.
It is also the false-alarm rate you agree to accept. If there is truly no difference, a test at α = 5% will still declare one in about 5% of cases, purely by chance. Lowering α to 1% makes false alarms rarer, but a real effect then needs more data to be detected; raising it to 10% catches more real effects and more false ones.
So the level is a business choice, not a law of nature. Use a stricter level when a false win is expensive or hard to undo, and a looser one for cheap, reversible changes. What matters most is when you choose it: before looking at the data. Picking α after seeing p = 0.07 ("let's use 10%") turns the threshold into decoration.
With many comparisons at once, each at 5%, the chance that at least one is a false alarm grows well above 5%, which is why multi-variant tests correct the level.
Example
CatChow compares two checkout versions: 120 of 3,000 orders against 150 of 3,000. The two-sided p-value is 0.062.
- At α = 5% the result is not significant: 0.062 > 0.05.
- At α = 10% it would be significant: 0.062 < 0.10.
The data are the same; only the line moved. That's why the team writes α = 5% into the test plan before launch, and doesn't reopen it after seeing 0.062.
Common mistakes
- Choosing α after seeing the p-value. The threshold only protects you if it was fixed in advance.
- Reading α as the chance the result is wrong. It is the false-alarm rate when there is no effect, not the probability that this particular win is false.
- Treating 0.05 as a cliff. p = 0.049 and p = 0.051 carry almost the same evidence.
- Keeping 5% for every comparison in a multi-variant test. The overall false-alarm rate climbs with each extra comparison.
Learn it in the course
- The t-test · Comparing the means of two groups and deciding whether the gap is bigger than chance.
- p-value and significance · What a p-value really says, what it doesn't, and statistical versus practical significance.
- Power and sample size · How many users an experiment needs to detect the effect you care about.
- Pitfalls: peeking, multiple testing, novelty effect · The most common ways real A/B tests go wrong, and how to avoid them.