Skip to content
Log in
← Statistics glossary

sampling bias

A systematic difference between sample and population caused by how observations were selected.

Sampling bias is a systematic difference between the people you measured and the people you want conclusions about, caused by how the sample was chosen. Unlike random noise, it doesn't average out: collecting more data the same way repeats the same error with more confidence.

It comes in a few familiar forms. Self-selection: people choose to answer, and those who care most (the happiest or the angriest) answer more. Non-response: invited people don't reply, and the ones who reply differ from the ones who don't. Coverage: part of the population can't be reached at all, for example customers without an email address. Survivorship: you only see the users who stayed, not the ones who left.

The margin of error doesn't include any of this. It assumes the answers are a random pick from the population. That is why a representative sample is about who answers, not how many.

What helps: invite a random list rather than whoever is around, send a reminder, report the response rate, and compare early with late answers (late answerers are often more like non-answerers).

Example

CatChow posts a link to a delivery survey in its customer community and gets 600 answers: 70% want evening delivery.

The community is where the most engaged customers spend time, and those who bothered to answer care about delivery. A random sample of 400 subscribers invited by email, with one reminder, finds 38%. The 600 community answers weren't wrong about the people who gave them; they just didn't stand for all subscribers. No margin of error could have warned about that.

Common mistakes

  • Fixing bias with volume. 10,000 self-selected answers are no more representative than 500.
  • Ignoring the response rate. At 5% response, the silent 95% decide how far off you might be.
  • Surveying only active users about churn. The people who left are exactly the ones missing.
  • Trusting a margin of error on a convenience sample. It describes random samples only.

Learn it in the course