one-tailed vs two-tailed test
A two-tailed test detects a difference in either direction; a one-tailed test only in a direction fixed in advance.
A two-tailed (two-sided) test asks whether B differs from A in either direction. A one-tailed (one-sided) test asks only one direction: is B higher? (or only: is B lower?).
The names come from the picture. Under "no difference" the test statistic follows a bell curve, and the results that count as significant sit in its tails. A two-tailed test at α = 5% splits them into 2.5% on each side; a one-tailed test puts all 5% on one side. So for the same data, when the result points the tested way, the one-tailed p-value is half the two-tailed one.
That halving is the danger. If you pick the one-tailed test after seeing which group is ahead, you get a smaller p-value for free and roughly double your chance of a false win. A one-tailed test is honest only when you chose it in advance and would act the same way on "no difference" as on "B is worse", for example a cheaper version you'll ship unless it is clearly better. It also can't flag a change in the other direction at all.
When in doubt, use two-tailed. It is the default in most tools for this reason.
Example
CatChow's checkout test: A converts 120 of 3,000 (4.0%), B 150 of 3,000 (5.0%). The z statistic is 1.87.
- Two-tailed: , not significant at 5%.
- One-tailed "is B higher?": , significant at 5%.
Same numbers, opposite verdicts. If the team had written "one-tailed, B higher" into the plan before launch, 0.031 stands. If they switched after seeing B ahead, the honest p-value is still 0.062.
Common mistakes
- Switching to one-tailed after seeing the data. It halves p without any new evidence.
- Using one-tailed when a drop would matter. A one-tailed test for "higher" can't flag that B hurts conversion.
- Comparing p-values from different test types. A one-tailed 0.03 and a two-tailed 0.06 are the same evidence.
- Not writing the choice down. If it isn't in the plan, readers will assume it was picked afterwards.