Interpreting the inferences a report makes
What you are being asked to do
- Reports do not only quote percentages. They quote confidence intervals for differences, average changes, and claims that something is "statistically significant".
- In this standard you interpret those inferences — you are not asked to carry out the bootstrapping or randomisation that produced them (that is AS91582 and AS91583).
- What you must be able to do:
- read a confidence interval for a difference and say what it licenses
- explain what "significant" does and does not mean
- spot a conclusion that goes beyond the inference reported.
Confidence intervals for a difference
- Studies often report a difference with an interval, for example:
"Patients on the new programme walked on average 340 m further at 12 weeks (95% CI: 120 m to 560 m)."
-
The reading:
- The best estimate of the difference is 340 m.
- Differences between 120 m and 560 m are consistent with the data.
- The method used produces intervals containing the true difference about 95% of the time.
-
The decisive question is whether the interval contains zero.
| Interval | What it means |
|---|---|
| 120 m to 560 m — entirely above zero | The programme group did better; the direction is established |
| −40 m to 300 m — contains zero | The data are consistent with no difference, and even with the programme being slightly worse |
| −600 m to −80 m — entirely below zero | The programme group did worse |
- An interval that contains zero does not prove the effect is absent. It means this study could not establish a direction.
Width tells you as much as position
- A narrow interval means a precise estimate; a wide one means the study was too small or too variable to pin the effect down.
- "CI: 120 m to 560 m" establishes that the programme helps, but leaves genuine doubt about whether the benefit is trivial or large — and those have very different practical implications.
- Always comment on width when the interval is wide. It is a criticism most students miss.
"Statistically significant"
- Statistically significant means only: the observed effect is larger than would readily be produced by chance alone.
- It does not mean:
- the effect is large — with a huge sample, a trivial difference becomes significant
- the effect is important — importance is a contextual judgement, not a statistical one
- the effect is causal — that depends entirely on the study design
- the result is certain.
- The reverse error matters too: "not significant" does not mean "no effect". A small study can easily fail to detect a real effect.
Practical versus statistical importance
- Ask how big the effect is in the units of the context, and whether that size would change any decision.
- A weight-loss programme producing a significant average loss of 0.4 kg over a year is statistically real and practically irrelevant.
- A 2% reduction in serious road injuries is a small percentage and a large number of people.
- The judgement requires contextual knowledge, which is exactly what Excellence in this standard asks for.
Relative and absolute change
- Reports routinely quote the relative change because it is the larger-sounding number.
"The new screening programme cut the risk of the disease by 50%."
- If the risk went from 4 in 10 000 to 2 in 10 000, then:
- the relative risk reduction is 50%
- the absolute risk reduction is 0.02 percentage points
- about 5 000 people must be screened to prevent one case.
- All three describe the same finding. A report that gives only the relative figure is presenting the most impressive true version, and saying so is a strong, well-defined criticism.
Inference is about the population, not the sample
- A report that says "in our sample, 62% improved" has made no inference at all — it has described the sample.
- An inference is a statement about the population, and it must carry uncertainty. Watch for reports that describe the sample and then conclude about the population without ever bridging the gap.
Worked ExampleInterpreting a reported inference
A report on a workplace wellbeing programme states:
"Employees who completed the programme took on average 1.2 fewer sick days over the following year than those who did not (95% CI: −0.3 to 2.7 days). This statistically significant result shows the programme cuts absenteeism by 1.2 days per employee. Rolled out across our 4 000 staff, that is 4 800 working days saved."
Evaluate the report's use of its own inference.
Step 1 — Check whether the interval contains zero
The interval runs from −0.3 to 2.7 days, and zero lies inside it.
This is the central error, and it invalidates everything built on top of it.
Step 2 — Interpret the interval properly
Step 2a — State what it licenses.
Step 2b — Comment on the width. The interval is 3 days wide around an estimate of 1.2 days — the uncertainty is more than twice the size of the estimate. This is characteristic of a study that was too small (or the outcome too variable) to answer the question it asked.
Step 3 — Examine the extrapolation
The report multiplies 1.2 days by 4 000 staff to reach 4 800 days saved.
Fault 1 — it uses the point estimate as though it were certain. Applying the interval instead gives a range from −1 200 days (a loss) to 10 800 days, which is not a usable planning figure. Multiplying a point estimate by a large number hides the uncertainty rather than removing it, and produces a precise-looking figure with no support.
Fault 2 — it generalises from completers to all staff. The estimate comes from employees who completed the programme; the 4 000 figure is the whole workforce. Rolling out a programme does not mean everyone completes it, and completers are a self-selected, more motivated group.
Step 4 — Examine the causal claim
The report says the programme "cuts" absenteeism.
Employees chose whether to complete the programme, so this is an observational comparison. Those who completed a voluntary wellbeing programme are plausibly healthier, more organised, or in less physically demanding roles to begin with — any of which would produce fewer sick days with no effect from the programme whatsoever.