sampling bias
A systematic difference between sample and population caused by how observations were selected.
Sampling bias is a systematic difference between the people you measured and the people you want conclusions about, caused by how the sample was chosen. Unlike random noise, it doesn't average out: collecting more data the same way repeats the same error with more confidence.
It comes in a few familiar forms. Self-selection: people choose to answer, and those who care most (the happiest or the angriest) answer more. Non-response: invited people don't reply, and the ones who reply differ from the ones who don't. Coverage: part of the population can't be reached at all, for example customers without an email address. Survivorship: you only see the users who stayed, not the ones who left.
The margin of error doesn't include any of this. It assumes the answers are a random pick from the population. That is why a representative sample is about who answers, not how many.
What helps: invite a random list rather than whoever is around, send a reminder, report the response rate, and compare early with late answers (late answerers are often more like non-answerers).
Example
CatChow posts a link to a delivery survey in its customer community and gets 600 answers: 70% want evening delivery.
The community is where the most engaged customers spend time, and those who bothered to answer care about delivery. A random sample of 400 subscribers invited by email, with one reminder, finds 38%. The 600 community answers weren't wrong about the people who gave them; they just didn't stand for all subscribers. No margin of error could have warned about that.
Common mistakes
- Fixing bias with volume. 10,000 self-selected answers are no more representative than 500.
- Ignoring the response rate. At 5% response, the silent 95% decide how far off you might be.
- Surveying only active users about churn. The people who left are exactly the ones missing.
- Trusting a margin of error on a convenience sample. It describes random samples only.
Learn it in the course
- Samples, populations and the normal distribution · Why we study samples to learn about all users, and why the bell curve shows up everywhere.