Surveys & Sample Sizes: Who Did They Ask?
Why 500 random people beat a million volunteers, how response bias haunts every NPS score, and the questions that separate research from theater.
What you'll learn
- Judge a survey by representativeness first and sample size second
- Spot the three ways surveys mislead: who was asked, who answered, and how the question was worded
- Treat small-sample findings as direction, not proof — without dismissing them entirely
A slide says “87% of customers want this feature.” Impressive — until you learn the survey went to attendees of the annual superfan conference, 15 of whom answered. The number is real. The claim it’s dressed up as (“customers want this”) is not. Surveys run the corporate world — NPS, engagement scores, market research, the pulse-check after every reorg — and they’re all vulnerable to the same three leaks. Learn the leaks and you can read any survey result in ten seconds.
The words that matter
The words that matter
- Population
- Everyone you actually want to know about — all customers, all employees, the whole market.
- Sample
- The subset you actually heard from. Every survey conclusion is a leap from sample to population — the leap is where the danger lives.
- Representative sample
- A sample that looks like the population in the ways that matter. Achieved by random selection, not by volume.
- Margin of error
- The honest wobble on a survey number. Around 1,000 random respondents gives roughly ±3 points; smaller samples wobble more.
- Selection bias
- The people you invited weren’t typical in the first place (conference superfans).
- Response bias
- The people who bothered to answer differ from those who didn’t — typically the delighted, the furious, and nobody in between.
- Leading question
- Wording that walks respondents toward an answer: “How much do you love the new dashboard?”
Random beats big
The most counterintuitive fact in all of survey-land: 500 randomly chosen respondents beat a million volunteers. A properly random sample of ~1,000 people estimates a whole country’s opinion to within about ±3 points — that’s why national polls work at all. Meanwhile the most famous polling disaster in history, the Literary Digest’s 1936 US-election poll, collected 2.4 million responses and still got the answer spectacularly wrong, because its mailing lists (car owners, telephone subscribers — luxuries at the time) skewed wealthy. Volume didn’t fix the skew. Volume never fixes skew; it just makes the wrong answer more precise.
The corporate translation: your NPS survey, your engagement survey, and your feature poll are almost never random samples. They’re whoever showed up — and who shows up follows a reliable pattern.
The population is mostly a quiet, moderately satisfied middle — but the survey inbox fills with its loudest edges.
Text description of this diagram
On the left sits a grid of thirty dots — the full customer population. Most are gray (quietly, moderately satisfied); a scattered few are green (delighted fans) or red (frustrated critics). A dashed arrow labeled who bothers to reply carries dots to the right-hand panel: the actual survey responses. That panel is almost entirely green and red dots — the delighted and the furious — with barely any gray. The amber banner explains the distortion: the quiet middle never fills in the form. The survey isn’t measuring the population; it’s measuring who had feelings strong enough to volunteer them.The three leaks — and false precision
Leak one: who was invited (selection bias). Superfan conferences, users of the beta, people who opened the email — each invitation list pre-filters reality. Leak two: who answered (response bias). Even a perfect invitation list leaks at the reply stage, because motivation to respond correlates with strong feelings. Leak three: how it was asked. “How much do you love the new dashboard?” isn’t a question; it’s a request for a compliment. Order matters too — ask about pay right after asking about layoffs and watch the scores move.
Then there’s the small-n problem, best caught by its tell: false precision. “83.3% of users prefer the new design” sounds authoritative until you realize it means five of six people in a hallway test. With samples that small, one changed mind swings the number by 17 points. Hallway tests are genuinely useful — for direction, for catching disasters, for deciding what to test properly. They’re just not evidence in the “greenlight the roadmap” sense. The polite phrasing when you spot it: “promising signal — what would it take to check it at a real sample size?” (And for that, module 6.)
Common misunderstanding
“A bigger response count means a more trustworthy survey.” Not if the sample is skewed — a million self-selected replies just measure the skew with impressive precision, which is exactly how the Literary Digest got 2.4 million responses and the wrong president. Representativeness first, size second. A small random sample honestly wobbles; a huge biased one confidently lies.Spot the leak
Try this at work
For any survey stat, ask the three-question audit: Who was invited? How many answered — out of how many? What was the exact wording? You’ll be amazed how often the deck doesn’t say — and how much the answer changes once it does. Bonus move for engagement-survey season: check the response rate before the scores; a 30% response rate means the scores describe a self-selected third of the team.The bottom line
A survey measures who answered, not who exists. Random and representative beats huge and biased, wording steers answers, and a percentage from a dozen people is a hint wearing a suit. “Who did they ask?” is the whole skill in four words.Why it matters
Surveys aren’t bad — they’re often the only affordable window into what people think, and courses like Customers & Market lean on them for personas and research. The skill is calibration: knowing whether a given number is a census, a decent estimate, or three enthusiastic conference-goers. Next, the strongest evidence machine in the toolkit — the A/B test (module 9) — and then the phrase it finally lets you decode, statistically significant (module 10).
Quick check
1. Which result should you trust more for "what do our customers think?"
2. NPS surveys tend to over-collect which voices?
3. "83.3% prefer the new design (n=6)." The right way to treat this:
Answers explained
- B is correct — a random cross-section represents the population; superfans represent superfans, no matter how many of them reply. (If you picked A: volume can’t cure skew — it just sharpens the wrong answer. If you picked C: picking the friendlier score is how survey theater works.)
- C is correct — responding takes effort, and strong emotion supplies it, so the quiet middle is systematically underweighted. (If you picked A: NPS’s scale doesn’t fix who chooses to answer. If you picked B: delighted fans are just as over-represented as critics.)
- A is correct — five of six people is a real signal worth following up, and nothing more; the decimal is costume jewelry. (If you picked B: one changed mind moves that number 17 points. If you picked C: small samples are fine for direction — dismissing them entirely throws away cheap, early learning.)