Collecting and processing data: validity and reliability
Validity and reliability are different things
-
These two words are used loosely in conversation and precisely in this standard. Confusing them costs marks at Merit and Excellence.
-
Validity — does the method measure what it claims to measure?
- Affected by: uncontrolled variables, an inappropriate range, an unsuitable measurement technique, sampling bias.
- An invalid method gives an answer to the wrong question.
-
Reliability — would the same method give consistent results if repeated?
- Affected by: number of trials, random variation, precision of instruments.
- An unreliable method gives inconsistent answers to the right question.
-
The crucial point: a method can be reliable but invalid. If a lamp warms the water while you increase light intensity, repeating the experiment ten times gives beautifully consistent results — consistently measuring the combined effect of light and temperature. Repetition cannot fix a validity problem. It only reduces random variation.
Reliability: repeats and means
- Do at least three trials at each level of the independent variable.
- Calculate the mean of the trials at each level. This reduces the effect of random variation, because random errors are as likely to be too high as too low and partly cancel out.
- Judge reliability by how close the repeats are to each other — the spread. Repeats clustered tightly indicate reliable data; widely scattered repeats do not, and the mean of scattered data is not trustworthy.
- Identify anomalies — results clearly outside the pattern of their repeats. State them, and say whether you excluded them from the mean and why. Never delete a result silently.
Two kinds of error
-
Distinguishing these is a Merit requirement, because the standard names "sources of errors" explicitly.
-
Random error — unpredictable variation, scattering results both above and below the true value.
- Causes: slight differences in timing, judgement in reading a scale, natural variation between organisms.
- Reduced by: more repeats and taking means. It cannot be eliminated, only averaged down.
-
Systematic error — a consistent bias shifting all results in the same direction.
- Causes: an instrument not zeroed, a scale reading consistently high, a stopwatch started late every time, a water bath running 2 °C above its setting.
- Not reduced by repeats at all — repeating gives the same wrong answer more precisely.
- Reduced by: calibrating instruments, checking against a known standard, improving technique.
-
This distinction matters practically: if your repeats agree closely but your results look wrong, suspect a systematic error. Consistency is evidence about random error only.
Sampling bias
-
Sampling bias occurs when the sample is not representative of what you claim to be describing. The standard names it explicitly at Merit.
-
It is the field equivalent of an uncontrolled variable: it produces a result that is wrong in a consistent direction and cannot be fixed by collecting more of the same biased sample.
-
Common sources:
- Convenience sampling — sampling where it is easy to reach, such as the edge of a site or the accessible side of a stream.
- Selecting "typical" or "good" specimens — which builds the investigator's expectation into the data.
- Sampling at one time of day when the organism's activity or the variable follows a daily rhythm.
-
Reduced by: random sampling (quadrat positions from random coordinates), systematic sampling (a transect at fixed intervals), adequate sample size, and sampling across the full range of the gradient rather than at one point.
Processing data
-
Raw data is what you record. Processed data is what you calculate from it, and the standard requires processing.
-
Typical processing:
- Means of repeated trials at each IV level.
- Rates, where a quantity is divided by time — often the most useful form in biology.
- Percentages or percentage change, which allows comparison between organisms of different starting size.
- A measure of spread — range or standard deviation — which shows how variable the repeats were and so supports a claim about reliability.
-
Presenting data:
- Tables should have clear headings with units, consistent decimal places, and raw and processed data distinguished.
- Graphs should have the IV on the x-axis and DV on the y-axis, labelled axes with units, a sensible scale, and plotted points. Add a line of best fit where a trend is claimed — and it need not be straight.
- Error bars showing the range or standard deviation are strong evidence of thinking about reliability, because they make the spread visible and show whether differences between levels are larger than the variation within them.
-
Graph type follows data type: a line graph for a continuous IV such as temperature or concentration; a bar graph for a categorical IV such as species or site.
Selective advantage
- Not applicable — this is a skills standard, assessing investigative process rather than biological content.
Worked Example
Worked Example
A student measures the volume of oxygen produced by catalase at five temperatures, three trials each.
| Temp (°C) | Trial 1 | Trial 2 | Trial 3 | Mean |
|---|---|---|---|---|
| 10 | 4.0 | 4.5 | 4.0 | 4.2 |
| 20 | 8.5 | 8.0 | 8.5 | 8.3 |
| 30 | 14.0 | 13.5 | 14.5 | 14.0 |
| 40 | 18.5 | 2.0 | 19.0 | 13.2 |
| 50 | 6.0 | 6.5 | 6.0 | 6.2 |
The student later finds the water bath was running 2 °C above its set temperature throughout.
Evaluate the reliability and validity of these data, and explain what should be done.
Answer:
Reliability — good, with one exception.
At 10, 20, 30 and 50 °C the three trials are close together, differing by at most 1.0 cm3. Tightly clustered repeats indicate that random error is small and the means at those temperatures can be trusted.
At 40 °C the trials are 4.0, 2.0 and 19.0 — trial 2 is wildly out of line with the other two, which agree closely with each other.
Handling the anomaly. Trial 2 at 40 °C is an anomaly. It should be identified and investigated, not silently deleted. Plausible causes include a gas syringe not sealed properly, a leak in the apparatus, or the potato discs not being added at the intended moment.
The mean of 13.2 is badly misleading: it is lower than the mean at 30 °C, which would suggest the rate falls between 30 and 40 °C. Both other trials at 40 °C are around 18.75, which fits the rising trend. So the anomaly, if left in, would change the shape of the relationship and lead to a false conclusion.
The correct action is to exclude trial 2 from the mean and state clearly that it was excluded and why, giving a mean of 18.75, and ideally to repeat that trial.
Validity — compromised by a systematic error.
The water bath running 2 °C above its setting throughout is a systematic error: it shifts every measurement in the same direction, so the true temperatures were 12, 22, 32, 42 and 52 °C rather than those recorded.
Three consequences follow:
- The repeats give no warning of it. Because the offset was consistent, the trials agree closely with each other. This is the key point: close agreement is evidence about random error only, and says nothing about systematic error.
- Repeating cannot fix it. More trials would simply reproduce the same offset more precisely.
- The shape of the relationship is unaffected, but every temperature value is wrong by 2 °C. So a conclusion about the pattern — rate rises to an optimum then falls — remains sound, while any conclusion about the specific optimum temperature is invalid.
What should be done.
- Correct the recorded values by adding 2 °C, or repeat the investigation with a calibrated water bath, checking its actual temperature with an independent thermometer.
- Calibrate instruments before starting in future, and check the bath temperature at the start and end of every trial rather than trusting the setting. This is precisely the sort of justified methodological choice that Excellence rewards.
- Repeat the 40 °C trial to replace the anomaly.
What this shows overall. The data are reliable but not fully valid — a combination students often miss, because consistency looks like quality. Reliability was demonstrated by the close repeats; validity failed because the instrument measured something other than what was recorded. Recognising that these are separate properties, and that one cannot substitute for the other, is what an evaluation at Excellence has to show.