The evaluation is where you turn on your own work. Not to apologise for it — to show you understand what your method could and could not deliver. The mark is not for finding faults; it is for tracing what each fault did to your result and suggesting a fix that would actually work in a school lab.
📘 What you need to know
Every point follows the same loop: weakness → impact → improvement.
Systematic errors shift every reading the same way. Random errors scatter readings unpredictably.
Repeats reduce random error; only better technique or calibration removes systematic error.
Weaknesses cause errors, limitations restrict what your conclusion covers, assumptions are simplifications you made.
Improvements must be relevant to the weakness and realistic for a school laboratory.
Never write “human error”. Name the specific step that went wrong.
Prioritise the one or two problems that mattered most to your result.
Start by evaluating your hypothesis
Your conclusion said whether the data supported the hypothesis. The evaluation asks a harder question: how strongly, given everything you now know about your data.
Even a supported hypothesis can rest on shaky ground. If the standard deviation at 50 °C was double that anywhere else, the exact point where denaturation sets in is less certain than the neat curve suggests — and saying so is a stronger piece of writing than claiming the graph settled it.
The two kinds of error
Telling these apart matters, because they do different damage and they need different fixes.
Look at the shape of the left-hand curve: it still peaks in the right place. A systematic error can leave your trend intact while making every value wrong.
Type
What it looks like in your data
Biology examples
How to reduce it
Systematic
Every reading off in the same direction; repeats agree with each other but not with reality
A water bath running 2 °C below its setting; a balance never tared; surface water left on tissue before weighing
Calibrate instruments; improve the technique; add a control
Random
Readings scatter above and below; large standard deviations
Natural variation between organisms; reaction time on a stopwatch; judging a colour change by eye
More repeats and a mean; a more objective measurement
The biology-specific one: natural variability between living samples is the single most common source of random error in a biology investigation, and it is worth discussing properly. Two leaves from the same plant are not identical, and no amount of careful technique changes that.
Weakness, limitation, assumption
These three get lumped together, and separating them is an easy way to make an evaluation look organised.
Limitations are not failings. Every investigation is limited — the skill is knowing exactly where your conclusion stops being safe.
Close the loop every time
This is the structure that separates a strong evaluation from a list of complaints. Each point needs all three parts, and the middle one is where most marks are won and lost.
“There may have been errors in the timing” is a weakness with no impact and no fix. It reads as filler, because it is.
WORKED EXAMPLE
In the amylase investigation the end point was judged by eye — when the iodine stopped turning blue-black. Evaluate this as a weakness.
Weakness
The end point was judged by eye, and the colour faded gradually rather than disappearing at one clear moment. This is a systematic error, because the same person tends to call the end point at the same slightly late stage every time.
Impact on the result
Every recorded time is slightly too long, so every calculated rate is slightly too low. The whole curve is shifted down, though the position of the optimum at 40 °C is unaffected because the shift applies to all temperatures. It also made the readings at 10 and 50 °C, where clearing was slowest, harder to judge and more scattered.
Improvement
Use a colorimeter with a fixed filter and stop timing at a set absorbance value, or record with a data logger so the end point is defined by a number rather than a judgement.
Named error type, direction of the shift, specific fixNotice the impact says which way the numbers moved. “It made the results less accurate” would say almost nothing.
WORKED EXAMPLE
Evaluate the variability of the enzyme and starch solutions as a source of random error.
Weakness
Small differences in mixing, in the time taken to transfer the tube into the water bath, and in how long solutions had stood before use, varied unpredictably between trials. This is a random error.
Impact on the result
It scattered the repeat times either side of the true value, producing the standard deviations of 1.4–4.0 s. The largest spread was at 50 °C, which widened the error bars exactly where the falling section of the curve needed to be read most carefully.
Improvement
Pre-warm the enzyme and substrate separately in the water bath for a fixed 10 minutes before mixing, start the timer at the moment of mixing, and raise the number of repeats from three to five so the mean is less affected by any one trial.
Random error, quantified from your own data, with a workable fixQuoting your own standard deviations is what makes this an evaluation of your experiment rather than a generic paragraph.
Limitations and assumptions
These are not errors. They are the boundaries around your conclusion, and stating them shows you know exactly how far your result reaches.
WORKED EXAMPLE
Give one limitation and one assumption for the amylase investigation, with an improvement or check for each.
Limitation
Only one source of amylase was tested, at a single pH, over a 40 °C range in 10 °C steps.
So: the conclusion that the optimum is 40 °C applies to this enzyme under these conditions, and cannot be generalised to amylase from other organisms. The 10 °C steps also mean the true optimum could lie anywhere between about 35 and 45 °C.
Extension: repeat with 2 °C steps between 30 and 50 °C to locate the optimum more precisely.
Assumption
It was assumed the buffer held the pH constant throughout, and that the tubes reached the set temperature before the enzyme was added.
Check: measure the pH of each tube at the start and end with a calibrated meter, and use a thermometer in a dummy tube to confirm the contents had reached temperature.
Scope stated, and a way to test what you assumedA limitation naturally suggests the next investigation, which is a good place for an evaluation to end.
Writing it well
Never write “human error”. It names nothing and fixes nothing. Say which step, done by whom, in what way.
Prioritise. Two problems examined properly beat eight listed. Choose the ones that most affected your final value.
Use your own numbers. Quote your standard deviations, your anomaly, your observations. A paragraph that would fit any experiment fits none.
Keep improvements realistic. A colorimeter, a data logger, more repeats and a randomised run order are all school-lab plausible. An electron microscope is not.
Do not fix what was not broken. Suggesting more repeats when your standard deviations were already tiny shows you did not read your own data.
A quick self-test: read your evaluation and see whether swapping in a different experiment’s title would still make it true. If it would, it is too generic to score.
💡 Exam tip
Structure every point as weakness, impact, improvement — and never skip the impact.
Say which way the error pushed your result: too high, too low, or more scattered.
Label each error as systematic or random, and match the fix to the type.
Separate weaknesses from limitations and assumptions, and say which you are discussing.
Quote your own standard deviations and observations as evidence.
Comment on how strongly your data supported the hypothesis, not just whether it did.
End with an extension: the investigation your limitations point towards.
⚠ Common mix-up
“Human error” as an explanation for anything.
Listing weaknesses with no impact on the result stated.
Suggesting more repeats for a systematic error. Repeats do not touch it.
Calling natural variation between organisms a systematic error. It is random.
Confusing a limitation with a weakness. Testing one species is a limitation, not a flaw.
Improvements nobody could carry out in a school laboratory.
A generic evaluation that never mentions a single number from your own data.
Apologising instead of analysing. The evaluation is a scientific judgement, not a confession.
That completes the inquiry cycle. Up next: Exploring a Problem — because a good evaluation always ends by pointing at the next investigation, and the cycle starts again.
Want this explained one-to-one?
Book a free session with an experienced IB Biology tutor and get your trickiest topics made simple.