Evaluating the research process
What this section is for, and what it is not
- The standard asks you to evaluate the research process — the strengths and weaknesses of how you did it, and how they affect the validity of your findings.
- It is not an apology. A list of things that went wrong, with no consequence attached, is worth very little.
- It is not about the results. We found high nitrate, which was disappointing is about the river, not about the method.
- Every point needs three parts: what you did, how it affected the data, and what it means for the confidence in your conclusion.
- At Excellence you must also discuss alternative research methods and their implications — not just what went wrong, but what you would do instead and what difference it would make.
The vocabulary that makes this section precise
- Validity — did you measure what you intended to measure? A design that confounds two variables has a validity problem.
- Reliability — would you get the same result if you did it again? Repeat visits test this directly.
- Accuracy — how close is a reading to the true value? Set by your equipment and technique.
- Bias — a systematic error that pushes results consistently in one direction. Excluding wet days is a bias, not random error.
- Representativeness — do your sites and times stand for the wider catchment and year? Five sites in March represent March at five points.
- Use these words correctly and the section writes itself; use them loosely and the whole evaluation reads as vague.
The strengths to claim, and how to evidence them
- Claim only strengths you can evidence. Our data was accurate is an assertion; the ranking of the five sites was identical on all four visits, and the between-visit spread at S1 was 1.0–1.4 mg/L against a between-site range of 7.7 mg/L is evidence of reliability.
- Controls are the easiest strengths to evidence: same time of day, same operator, same equipment, same sampling position, same antecedent conditions.
- A control site is a strength worth naming. Without S1 in the forest, no comparison with land use is possible at all.
- Repeat measurement is the strongest evidence you can offer, because it tests reliability directly rather than asserting it.
- Say what each strength protects. Holding time of day constant removes the daily cycle as a competing explanation of the spatial pattern.
The weaknesses to identify, in order of importance
- Design weaknesses — confounded variables, missing controls, a sample that cannot answer the aim. These matter most, because no amount of care fixes them.
- Sampling weaknesses — too few sites, too few visits, sites that are not representative, a deliberately excluded condition.
- Measurement weaknesses — equipment error, technique, one operator's judgement.
- Recording and processing weaknesses — transcription errors, lost sheets, inconsistent rounding.
- Rank them. A report that treats a transcription slip and a confounded design as equally important has not evaluated anything.
- Quantify where you can. Flow measurement error of about ±20 per cent, which is larger than the 6 per cent difference between S4 and S5 is a weakness with a consequence attached.
Alternatives, which is what Excellence requires
- The standard's Excellence criterion names it: discuss alternative research methods and their implications.
- For each alternative, give three things: what it is, what problem it fixes, and what it costs — in time, access, equipment or comparability.
- In the invented study: paired tributary sampling fixes the confound but needs landowner access; wet-weather visits bound the effect from above but reduce comparability and raise safety issues; a continuous logger settles whether four visits are representative but trades spatial coverage for temporal detail.
- Then choose. Say which alternative you would adopt first and why, and base the choice on which limitation is actually binding.
- The commonest error is to propose improvements to the part that already works — usually more precise equipment — because it is the easiest to write about.
Worked Example
Worked example
Evaluate the research process used in the invented Rerenga study, and discuss alternative methods.
Answer:
Step 1 — state the strengths with evidence, briefly.
Reliability was good and I can show it: the site ranking was identical on all four visits, and the between-visit spread at S1 (1.0–1.4 mg/L) was small against a between-site range of 7.7 mg/L. Holding time of day (9 am), operator, equipment and sampling position constant removed four competing explanations for the spatial pattern. The control site at S1 is what makes any land-use comparison possible at all.
Step 2 — rank the weaknesses, most serious first, each with its consequence.
- Design: distance and land use are confounded. Forest is upstream, pasture in the middle, township at the bottom. Consequence: my data cannot distinguish "land use raises nitrate" from "nitrate accumulates downstream". This limits what my conclusion can claim, not how precisely it can claim it.
- Sampling bias: wet days excluded. Sampling at least three days after rain made visits comparable and removed the conditions when pastoral runoff is greatest. Consequence: my land-use effect is a lower bound, not a central estimate.
- Measurement: flow error about ±20 per cent. Consequence: the 6 per cent fall between S4 and S5 is uninterpretable — dilution and in-stream processing cannot be separated. I have not claimed recovery.
- Representativeness: four visits in one month. Consequence: nothing in my study speaks to seasonal variation, and I have not implied otherwise.
Step 3 — discuss three alternatives, with what each fixes and what it costs.
- Paired tributary sampling — sample a forested and a pastoral tributary of similar size entering the same reach. Fixes: the confound, by holding distance roughly constant while land use varies. Costs: landowner access, which I did not have for two of the tributaries.
- Two wet-weather visits — sample within 24 hours of rain. Fixes: the lower-bound problem, by bounding the effect from above. Costs: comparability between visits, and a genuine safety limit in high flow.
- A conductivity or nitrate logger at S4 and S5 for a fortnight — continuous rather than snapshot data. Fixes: whether four visits are representative, and possibly the S5 question. Costs: equipment I do not have, and it would reduce the number of sites.
Step 4 — choose, and justify the choice against the ranking.
I would adopt paired tributary sampling first. My data show that precision is not the binding constraint: four visits agree closely, so random error is small. The binding constraint is a design confound, and a confound is not improved by better instruments or more repeats. Choosing the logger instead would produce a more precise answer to a question my design cannot ask.