Making the inference and writing the conclusion
What kind of claim this standard asks for
- At Level 2 the conclusion is a suggestive inference: the treatment appears to have caused a change in the response.
- The word "appears" is doing real work. It is stronger than "there is a difference" and weaker than "the treatment definitely works" — and it is exactly the level of confidence one experiment supports.
- You do not need a formal statistical test here. The evidence for your inference is the displays and measures you have already discussed: the shift, the difference in centres, and how much the groups overlap.
- A formal test of whether chance alone could explain the difference — the re-randomisation method — belongs to Level 3 (AS91583). It is not required, or expected, in this standard.
- What earns the marks instead is linking the evidence you have to a claim of the right strength, and saying honestly who that claim covers.
Why an experiment can support a causal claim at all
- Random allocation is the licence. Because the units were assigned to the two groups at random, the groups started out similar on average in every other respect — including things you never thought to measure.
- So the treatment is the only systematic difference between them. If the response differs, the treatment is the most plausible explanation.
- This is what an observational study can never do. If people choose their own group, the groups differ in many ways at once, and any difference in outcome could belong to any of them.
- The word "random" must appear in your justification. A conclusion that claims cause without pointing at the random allocation has not justified anything.
Weighing the evidence before you claim
- A difference in medians alone is not the whole story. Ask these questions before deciding how strongly to word the inference:
- How big is the shift compared with the spread? A shift much larger than the IQRs is convincing; a shift much smaller than them is not.
- How much do the groups overlap? Little overlap supports a firmer claim; heavy overlap calls for a cautious one.
- How many units were there? A difference found across 40 plants is more convincing than the same difference across 6.
- Match the strength of your wording to the strength of your evidence. "Clearly appears to have increased" and "appears to have slightly increased, though the groups overlap considerably" are both legitimate — using the first when the data only supports the second is not.
Worked ExampleWriting the inference from the displays
In a fertiliser experiment, 24 tomato plants were randomly allocated to two groups of 12. The treatment group received fertiliser; the control group did not.
Treatment median yield: 15 kg (box 12–18 kg). Control median: 11 kg (box 9–14 kg).
Write the inference and conclusion.
Step 1 — State the difference in centres, with context and units
Step 2 — Weigh it against the spread and the overlap
Each group's box spans about 5-6 kg, so a 4 kg shift is large relative to the spread — but the boxes do overlap between 12 and 14 kg, so some unfertilised plants out-yielded some fertilised ones. The shift is convincing; the separation is not complete.
Step 3 — Make the inference, at the right strength
Step 4 — Justify the causal wording
Step 5 — State who the conclusion covers
What a complete conclusion contains
- The inference — the treatment appears to have caused a change, and in which direction.
- The evidence — the difference in centres with numbers and units, plus a comment on spread and overlap.
- The justification — random allocation is why a causal claim is allowed.
- The limits — who the claim covers, and what would be needed to go further.
Honest limits
- The claim covers the experimental units you used. Extending it to a wider population needs those units to have been a random sample of that population as well as randomly allocated.
- One experiment is never proof. It is evidence, and the honest word for it is "appears to".
- A small or unclear difference is a legitimate result. Say that the experiment did not show a clear effect, and note that a larger number of units might reveal a smaller real effect that this one could not detect.
- Name anything that went wrong — a plant that died, a measurement taken a day late, a unit that got the wrong treatment — and say which way it would have pushed the result. Reporting this honestly gains credit; hiding it is the one thing that cannot be recovered from.