Skip to content
Log in
← Statistics glossary

population

The whole group we want conclusions about, e.g. all current and future visitors of the shop.

The population is everyone your conclusion is about: all active subscribers, all visitors of the checkout page this quarter, all users who signed up in March. The sample is the part of it you actually observe.

Naming the population precisely is the first step of any survey or experiment, because it decides who may be invited and what the result may be generalised to. "Our users" is not a population; "customers who ordered at least once in the last 90 days" is.

Its size matters less than people expect. For a large population, the precision of an estimate depends on the number of answers, not on the share of the population you reached: 385 answers give about ±5 points whether there are 50,000 customers or 5 million. Only when the population is small, and you reach a noticeable part of it, does its size lower the number you need (the finite population correction).

In experiments the population is often not a fixed list but a stream of future visitors. The sample you test on this month stands in for the people who will arrive later, which only works if this month is typical.

Example

CatChow wants to know how many subscribers would switch to evening delivery. The population is "the 2,000 subscribers with an active plan today", not "everyone who ever ordered".

For ±5 points at 95%, a huge population would need 385 answers. Because this one has only 2,000 people, the finite population correction lowers it to 323. If the population had 1,000,000 people, the answer would be 384, barely different from 385.

Common mistakes

  • Surveying one group and reporting on another. Answers from newsletter readers describe newsletter readers, not all customers.
  • Thinking a big population needs a big sample. Above roughly 100,000 people the required number of answers barely grows.
  • Leaving the population vague. Without a clear definition, you can't tell who was missing from the sample.
  • Assuming the test month stands for every month. Seasonality and campaigns change who arrives.

Learn it in the course

  • Samples, populations and the normal distribution · Why we study samples to learn about all users, and why the bell curve shows up everywhere.
  • Spread: range, SD and quartiles · Two couriers with the same average delivery time, and why only one keeps the promise: range, variance, standard deviation, quartiles, IQR and the coefficient of variation.
  • Outliers and skewed data · Why product metrics have long tails, how to find outliers with the IQR rule, and what to do with them: fix errors, segment different customers, and never delete whales without thinking.
  • Confidence intervals · Reporting a range of plausible values instead of a single number.
  • ANOVA: comparing many groups · Testing three or more variants at once without inflating false positives.