Sampling, representativeness and bias
From sample to claim
- Almost every report measures a sample and then makes a claim about a whole population.
- That claim is only trustworthy if the sample is representative — if it looks like the population in the ways that matter.
Start with the population, the variables and the measures
The standard lists population, measures and variables as the first features of a survey to evaluate. Get these straight before judging anything else.
- What population does the report claim to describe? "New Zealanders", "Auckland renters", "Year 12 students at this school" — the claim can only be as wide as the population actually sampled.
- What variable was measured? Not the headline word, the actual question asked.
- Does the measure match the claim? This is where reports most often break down:
- a headline about "stress" built on a question asking how many hours people slept
- a claim about "exercise" measured by asking whether people own a gym membership
- a claim about household income measured by asking one person to estimate it
- A measure that does not capture the variable makes everything downstream worthless, however large the sample or careful the sampling.
How was the data collected?
Survey method is a named feature, and it shapes who ends up answering.
- Face-to-face — high response rate, but people are less honest about sensitive topics with an interviewer present.
- Telephone — misses anyone without a phone or unwilling to answer unknown numbers; landline-only polls skew old.
- Online — cheap and fast, but excludes people with poor internet access and is easy to answer dishonestly or repeatedly.
- Postal — slow, with typically low response, and the people who bother to reply differ from those who do not.
- Say how the method affects this particular report. "It was an online survey, so students without home internet are underrepresented — and they are exactly the group the question about digital homework is about" is worth far more than listing the methods.
How was the sample chosen?
- Random sample — everyone in the population has an equal chance of being picked. This is the gold standard: it avoids systematic bias and lets the sample stand in for the population.
- Convenience sample — whoever is easy to reach (the first 30 people outside a shop). Usually not representative.
- Self-selected (voluntary response) — people choose to take part (a website poll, a text-in). Strongly biased: only those with strong opinions or time respond.
Sources of bias (non-sampling errors)
- Self-selection / selection bias — the way people enter the sample skews who is in it.
- Non-response bias — those who do not reply may differ from those who do, so the responses are lop-sided.
- Question wording — a leading or loaded question pushes people toward an answer ("Don't you agree that…?").
- Measurement bias — the way the response is collected distorts it.
Sample size
- A larger sample gives a more reliable result — but a big biased sample is still biased. Size does not fix a bad sampling method.
The questions to ask of any report
- Who was in the sample, and how were they chosen?
- How many were sampled?
- Who was left out, and could that skew the result?
Worked ExampleEvaluating a website poll
A news website runs an online poll: "Should the council build more cycleways?" Of 4,000 people who clicked, 78% said yes. The article concludes "most residents support more cycleways." Evaluate this.
Step 1 — Identify the sampling method
The sample is self-selected (voluntary response) — only people who visited the site and chose to click took part.
Step 2 — Explain the bias
People who click on such a poll tend to hold strong views on the topic, and website visitors are not a cross-section of all residents. So the 78% reflects the opinions of a keen minority, not the population — this is selection bias.
Step 3 — Effect on the conclusion
The conclusion "most residents support it" is not justified: the sample cannot be generalised to all residents, no matter that 4,000 responded.