Experiments, observational studies and what each can claim
The one distinction that decides everything
- In an experiment, the researcher decides who receives the treatment.
- In an observational study, the participants' own circumstances or choices decide.
- That single difference determines whether the report may say "causes" or only "is associated with".
The principles of experimental design
- Comparison — there must be at least two groups. A single group measured before and after has no comparison, so any change may be due to time, to the season, or to being studied.
- Random allocation — participants are assigned to groups by chance. This is the crucial step: it makes the groups similar on every variable at once, including variables nobody thought of or could measure.
- Replication — enough units in each group that the comparison is not swamped by natural variation.
- Control of other variables — everything except the treatment is kept as similar as possible between groups.
- Blinding where possible — participants (and ideally assessors) do not know which group they are in, removing the placebo effect and assessment bias.
Why random allocation is what licenses "causes"
- Suppose people who exercise more also sleep better, eat differently and have more leisure time. In an observational comparison these travel together, and the effect of exercise cannot be separated from them.
- If you randomly allocate people to an exercise programme, the two groups end up with similar distributions of sleep habits, diet and leisure time — not because you measured or matched them, but because chance distributes them evenly.
- So if the groups then differ in the outcome, the only systematic difference between them was the treatment.
- This is the entire argument, and it is worth being able to state in one sentence: random allocation balances all other variables between the groups, so a difference in outcome can be attributed to the treatment.
Random allocation versus random selection
- These are different, they do different jobs, and confusing them is a common and costly error.
| What it is | What it buys | |
|---|---|---|
| Random selection | Choosing who takes part, from the population | Generalisability — results extend to the population |
| Random allocation | Choosing which group each participant goes into | Causation — a difference can be attributed to the treatment |
- A study can have one, both or neither:
- Both: a randomly selected national sample, randomly allocated. Causal conclusions about the population — rare and expensive.
- Allocation only: volunteers randomly allocated. Causal conclusion about people like the volunteers. This is most medical trials.
- Selection only: a random national survey with no treatment. Population estimates and associations, no causation.
- Neither: a convenience sample of self-selected groups. Neither causation nor generalisation.
Types of observational study
- Cross-sectional — measured at one point in time. Cheap; cannot establish which came first.
- Case-control — start with people who have the outcome and compare with those who do not, looking backwards. Efficient for rare outcomes; vulnerable to recall bias.
- Longitudinal (cohort) — follow a group forward over time. Establishes the time order of events, which removes reverse causation but not confounding. New Zealand has two internationally known examples in the Dunedin and Christchurch longitudinal studies.
Why observational studies are still done
- This matters for a balanced evaluation. Do not dismiss an observational study merely for being observational — say why the design was chosen and what it can still support.
- Ethics: you cannot randomly allocate people to smoke, to be injured, or to live in poor housing.
- Practicality: you cannot randomly allocate a region's weather, or a family's income.
- Time and cost: effects that take 30 years to appear cannot wait for a trial.
- An observational study can support: an association, its size and direction, a hypothesis for further work, and — with enough consistency across studies, a plausible mechanism and a dose-response pattern — a strong causal case, even without a single experiment.
Worked ExampleClassifying a study and fixing its conclusion
Two reports appear in the same week.
Report 1: "We recruited 240 volunteers with poor sleep and randomly assigned half to a four-week evening screen-reduction programme and half to a waiting list. After four weeks, the programme group slept on average 34 minutes longer per night (95% CI: 12 to 56 minutes). Reducing evening screen use improves sleep."
Report 2: "In a survey of 5 000 randomly selected New Zealand adults, those who reported more than three hours of evening screen use slept on average 41 minutes less per night than those reporting less than one hour. Evening screen use is damaging New Zealanders' sleep."
For each report, classify the study, state what its conclusion may legitimately claim, and identify its main limitation.
Report 1 — classify
Step 1 — Who decided the groups? The researchers did, at random. This is a randomised experiment.
Step 2 — Check the design principles.
- Comparison: yes — a waiting-list group.
- Random allocation: yes.
- Replication: 120 per group, adequate.
- Blinding: not possible — participants necessarily know whether they are reducing screen use. This leaves a placebo/expectation effect and, since sleep was self-reported, possible reporting bias.
Step 3 — What may it claim?
Step 4 — Main limitation.
Report 2 — classify
Step 1 — Who decided the groups? The participants themselves, by how they live. This is an observational, cross-sectional study.
Step 2 — What may it claim?
Step 3 — Main limitation — and be specific.
The word "damaging" asserts causation, which this design cannot support, for two distinct reasons:
- Confounding. People with heavy evening screen use differ in other ways — shift work, longer commutes, younger age, higher stress, irregular schedules. Any of these independently reduces sleep, and the study cannot separate them from screen use.
- Reverse causation. Because everything was measured at one point in time, the direction is undetermined. People who cannot sleep may reach for a screen precisely because they are lying awake. The same data would appear if poor sleep caused screen use rather than the other way round.