Skip to content
Log in
← Statistics glossary

one-tailed vs two-tailed test

A two-tailed test detects a difference in either direction; a one-tailed test only in a direction fixed in advance.

A two-tailed (two-sided) test asks whether B differs from A in either direction. A one-tailed (one-sided) test asks only one direction: is B higher? (or only: is B lower?).

The names come from the picture. Under "no difference" the test statistic follows a bell curve, and the results that count as significant sit in its tails. A two-tailed test at α = 5% splits them into 2.5% on each side; a one-tailed test puts all 5% on one side. So for the same data, when the result points the tested way, the one-tailed p-value is half the two-tailed one.

That halving is the danger. If you pick the one-tailed test after seeing which group is ahead, you get a smaller p-value for free and roughly double your chance of a false win. A one-tailed test is honest only when you chose it in advance and would act the same way on "no difference" as on "B is worse", for example a cheaper version you'll ship unless it is clearly better. It also can't flag a change in the other direction at all.

When in doubt, use two-tailed. It is the default in most tools for this reason.

Example

CatChow's checkout test: A converts 120 of 3,000 (4.0%), B 150 of 3,000 (5.0%). The z statistic is 1.87.

  • Two-tailed: p=P(∣Z∣≥1.87)≈0.062p = P(|Z| \ge 1.87) \approx 0.062, not significant at 5%.
  • One-tailed "is B higher?": p=P(Z≥1.87)≈0.031p = P(Z \ge 1.87) \approx 0.031, significant at 5%.

Same numbers, opposite verdicts. If the team had written "one-tailed, B higher" into the plan before launch, 0.031 stands. If they switched after seeing B ahead, the honest p-value is still 0.062.

Common mistakes

  • Switching to one-tailed after seeing the data. It halves p without any new evidence.
  • Using one-tailed when a drop would matter. A one-tailed test for "higher" can't flag that B hurts conversion.
  • Comparing p-values from different test types. A one-tailed 0.03 and a two-tailed 0.06 are the same evidence.
  • Not writing the choice down. If it isn't in the plan, readers will assume it was picked afterwards.