Designing a fair experiment
The principles of a good design
- Random allocation — split the units into groups by chance (draw names, use random numbers on a calculator or spreadsheet).
- This is the single most important feature of an experiment. It balances out confounders — known and unknown — across the two groups.
- Random does not mean casual. "I put the keen students in group A" or "the first 15 on the roll" is not random, and it destroys the causal claim.
- A control group — a group that gets no treatment, or a fake one, giving a baseline to compare against.
- Without a control you cannot tell what the treatment did. If everyone improves, was it the treatment, or would they have improved anyway with practice, time or attention?
- Replication — enough units in each group.
- A difference based on two plants is far weaker than one based on twenty, because a couple of unusual units can swing a small group's centre entirely.
- More units also let smaller real effects show through the natural variation between individuals.
- Keep everything else the same — a fair test changes only the treatment; water, light, timing, instructions and measurement are held identical.
What can go wrong: confounding
- A confounding variable is something that differs between the groups as well as the treatment, giving a second possible explanation for any difference.
- Examples that catch students out:
- the treatment plants sat on the sunnier windowsill
- the treatment group was tested in the morning and the control after lunch
- the treatment group knew they were being studied, and tried harder
- A confounder is fatal to the conclusion, because you can no longer say which of the two differences caused the change.
- Random allocation handles the confounders you never thought of, which is exactly why it matters. Controlling variables by hand only protects you from the ones you can name.
Reducing bias
- Blinding — where possible the units, and the person measuring, should not know which group is which, so expectations cannot sway the result.
- Single-blind: the participants do not know.
- Double-blind: neither the participants nor the person measuring knows. This is the stronger version, because it also removes unconscious bias in the marking or measuring.
- A placebo — a fake treatment that keeps the control group's experience identical apart from the active ingredient. It matters because believing you have been treated can change the outcome on its own.
- Standardise the measurement — measure the response the same way, with the same instrument, at the same point, for every unit.
- Say what you could not blind, and why. Some experiments cannot be blinded — a student obviously knows whether music is playing — and identifying that honestly is worth more than pretending it was controlled.
Writing the plan
- State the conjecture — what you think the treatment will do, and why.
- Name the experimental units, the treatment and the control.
- Say exactly how the random allocation will be done — the method, not just the word "randomly".
- Name the response variable and how it will be measured, including the unit and the precision.
- List the variables you will hold constant, and say how.
- State how many units go in each group.
- Write the plan before collecting anything, and record any change you had to make once it was running — that record is itself worth credit.
Worked ExamplePlanning a fair experiment
Design an experiment to test whether a "study playlist" improves memory-test scores for a class of 30 students.
Step 1 — Conjecture and variables
- Conjecture: listening to the playlist while revising improves memory-test scores.
- Explanatory variable: playlist while revising, or silence.
- Response variable: score out of 20 on the memory test.
Step 2 — Random allocation
Randomly allocate the 30 students to two groups of 15 — e.g. draw names from a hat, or number them and use random numbers. Group A revises with the playlist; Group B revises in silence.
Step 3 — Keep it fair
Both groups revise the same material, for the same time, in the same room conditions, and sit the same test, marked the same way. The only difference is the playlist.
Step 4 — Replication
15 students per group is enough to see a difference beyond one or two unusual results.