Skip to content
Log in

Statistical Significance Calculator

Check whether the difference between two groups is bigger than chance, and see what the p-value really says.

Reading an A/B experiment? The A/B test results calculator focuses on lift and the interval.

Want to understand what a p-value is, step by step? The free lesson builds it from scratch. Learn the p-value and significance

Your two groups

Total: people, emails or visitors in the group. Successes: conversions, clicks or yes answers. Count each person once.

Group A

Rate: 4%

Group B

Rate: 5.5%

Test

Choose before you look at the numbers. Switching to one-sided after seeing the result halves the p-value and makes false wins twice as likely.

Advanced settings
Significance level (α)

5% is the usual choice. Pick it before you look at the data.

What to enter in these fields
Total
Everyone in the group who could have had the outcome: people, emails sent or visitors, not page views. A whole number from 1 to 1,000,000,000.
Successes
How many of them had the outcome you count: a conversion, a click, a yes answer. Count each person once. A whole number from 0 up to the total.
Test
Two-sided asks whether B differs from A in either direction and is the safe default. One-sided asks only whether B is higher (or only whether it is lower). Use it only if you decided so in advance and would act the same way on "no difference" and on a difference in the other direction.

Result

Enter the counts of both groups and press Check significance. The example is ready to try.

How to read the result

The verdict compares the p-value with the significance level you chose. Below it, the difference is called statistically significant: a gap this large would be unusual if the two groups really had the same rate.

The sentence under the verdict says what the p-value measures in your own numbers, and the picture shows it: the hatched tails are the results that would count as significant, the filled area is how much of the curve lies at or beyond your result.

Then look at the size of the difference and its interval in the details. A tiny difference can be significant with enough data, and a large one can miss significance in a small sample. Significance tells you whether to trust the direction, not whether the difference matters.

What statistical significance means

Group A converted 120 of 3,000 (4.0%) and group B 165 of 3,000 (5.5%). Even if both groups had the same true rate, two samples would rarely come out exactly equal, so the question is whether a gap of 1.5 points is ordinary noise.

The calculator measures the gap in units of its expected noise: z = 2.73, almost three times the typical random gap for samples this size. A gap this large or larger in either direction would appear by chance in only about 0.63% of such comparisons (p = 0.0063).

Because 0.0063 is below the 5% level, the difference is statistically significant. That is a statement about noise, not about causes: if the groups differ in more than one way, the data can't say which of those differences produced the gap.

One-sided or two-sided?

Take A at 120 of 3,000 and B at 150 of 3,000. The two-sided test gives p = 0.062: not significant at 5%. The one-sided test "is B higher?" on the same numbers gives p = 0.031: significant. Same data, opposite verdicts; the only difference is the question asked.

That is why the test has to be chosen before you look. If you pick the side after seeing which group is ahead, you halve the p-value for free and roughly double the chance of a false win. Textbooks also call the two options one-tailed and two-tailed tests.

A one-sided test is defensible when you decided on it in advance and would act the same way on "no difference" and on "B is worse", for example when B is a cheaper version you'll ship unless it's clearly better. When in doubt, use two-sided.

Method and formulas

The calculator uses the pooled two-proportion z-test, the same engine as the A/B test results calculator, so the two-sided numbers of both tools are identical for the same counts. xA and xB are the successes, nA and nB the totals:

pA = xA / nA,  pB = xB / nB,  d = pB − pA
pPool = (xA + xB) / (nA + nB)
SE0 = √(pPool (1 − pPool) (1/nA + 1/nB)),  z = d / SE0
two-sided:        p = P(|Z| ≥ |z|)
one-sided higher: p = P(Z ≥ z)
one-sided lower:  p = P(Z ≤ z)
significant = p < α

The two-sided interval is Newcombe's hybrid score interval for B − A at 1 − α. A one-sided test reports one bound at 1 − α: the lower end (for "B higher") or the upper end (for "B lower") of the Newcombe interval at 1 − 2α.

Each group needs at least 10 successes and 10 non-successes. Below that the normal approximation is unreliable, so the calculator shows the rates without a p-value rather than switching silently to another test.

Questions and answers

What does p = 0.03 mean in plain words?
If there were truly no difference between the groups, results at least as far apart as yours would turn up in about 3% of studies like this, by chance alone. It is not a 3% chance that there is no difference, and not a 97% chance that B is better.
Is 5% the right significance level? When would I use 1% or 10%?
5% is a convention, not a law. Use 1% when a false win is expensive, for example a change that is hard to undo. Use 10% for cheap, reversible decisions where missing a real effect costs more. Pick the level before the test and don't change it after seeing the p-value.
One-sided or two-sided: which should I pick?
Two-sided, unless you decided before collecting data that only one direction matters and that you'd treat "no difference" and "worse" the same way. Switching to one-sided after seeing the numbers turns a p-value of 0.06 into 0.03 without any new evidence.
The result is significant. Does that mean the difference is big or worth acting on?
No. Significance means the difference is unlikely to be pure noise. With a large sample, a gap of 0.1 percentage points can be significant and still not worth the effort. Look at the difference, its interval and what it's worth to the business.
The result is not significant. Are the two groups the same?
Not necessarily. The data may simply be too few to separate a real difference of this size from noise. The interval shows the range of differences that are still plausible; if it includes differences you care about, collect more data before concluding anything.
Can I use this for two customer segments rather than an experiment?
Yes, to ask whether the gap between them is bigger than noise. No, to ask why they differ: segments differ in many ways at once, so a significant gap between, say, mobile and desktop users doesn't show that the device caused it.
Can I check the numbers every day until they become significant?
No. Every look is another chance for noise to cross the line, so stopping at the first significant day makes false wins far more likely than the level you chose. Decide the sample size and when to look in advance; the post on peeking below explains why.
Why does the calculator refuse very small counts?
The z-test relies on an approximation that breaks down when a group has fewer than about 10 successes or 10 non-successes. Rather than show a p-value that could be badly off, the calculator shows the rates only. Collect more data, or use an exact test in a statistics package.

Learn more

Planning the next comparison?

To plan how many people you need so that a real difference shows up as significant, use the A/B test sample size calculator.

Comparing answers in a survey? Plan the number of answers with the survey sample size calculator.

Terms used here