IB Chemistry SLTopic 8 — Concluding and EvaluatingInternal assessmentPractical skill~14 min read
Evaluating
The evaluation is where students either pick up a lot of marks or write half a page of apologies. The difference is small: an apology says something went wrong, while an evaluation says what went wrong, which way it pushed the answer, and what you would change.
📚 What you need to know
Comment on how strongly your data supports the hypothesis, not just whether it does.
Tell systematic errors apart from random ones. They behave differently and need different fixes.
Repeating trials reduces random error. It does nothing to a systematic error.
Every weakness needs all three parts: weakness → impact → improvement.
Say which direction each error pushed your result. That is the sentence examiners look for.
Weaknesses, limitations and assumptions are three different things.
Improvements must be realistic in a school lab, and you should lead with the biggest problem first.
Systematic or random?
Almost everything in an evaluation depends on getting this split right, and there is a simple test. Ask what would happen if you ran the experiment twenty more times. If the results would spread out around the true value, the error is random. If all twenty would land on the same side, it is systematic.
This is why “I would repeat the experiment more times” is such a weak improvement. It only helps with the picture on the left, and most IA problems are on the right.
Type
Typical causes in a school lab
What actually reduces it
Systematic
Heat escaping a calorimeter; a balance that was never zeroed; a burette rinsed with water instead of the solution; reading a scale from the same wrong angle every time.
Change the apparatus or the procedure. Calibrate. Insulate. Rinse properly.
Random
Judging an endpoint by eye; reaction time on a stopwatch; small fluctuations in room temperature; the last drop hanging on the burette tip.
Repeat and average. Use an instrument instead of a judgement, such as a colorimeter or a pH probe.
One error can be both, depending on how it happens. If you always read the meniscus from slightly above, that is systematic. If your eye height wanders from reading to reading, it is random. When you write about parallax, say which one you mean and why — that single clarification shows you understand the difference rather than having memorised two lists.
Weakness, impact, improvement
Every point you make in an evaluation has three parts, and the middle one is where the marks sit. Students are good at naming problems and good at suggesting fixes. Very few say what the problem actually did to the number.
Test your impact sentence by asking whether it contains a direction. “This affected my results” has none. “This made the temperature rise too small” has one.
Weaknesses, limitations and assumptions
These three get lumped together and they are not the same. Being able to separate them gives you three distinct things to write about instead of one repeated one.
Term
What it means
Example
Weakness
Something in your method that caused error and could have been done better
Using a glass beaker rather than an insulated cup, so heat leaked away through the walls
Limitation
A boundary on what your conclusion is allowed to cover, even if nothing went wrong
Only concentrations from 0.50 to 2.50 mol dm−3 were tested, so nothing can be claimed above that
Assumption
A simplification you made in the calculation that is not exactly true
Taking the density and specific heat capacity of the solution as those of pure water
Limitations are the easiest marks in this section and most students skip them. If you tested five temperatures between 20 and 60 °C, then your trend is only established between 20 and 60 °C. Saying so is not weakness, it is precision about what you have shown — and it sets up a genuine extension: test a wider range and find out whether the pattern holds.
Lead with the biggest problem
An evaluation is not a list of everything imperfect about your afternoon. Two or three points, properly developed, beat eight one-liners. And you already know which points matter, because you worked out the percentage uncertainties when you processed the data.
Two of these bars are so short you can barely see them. Spending a paragraph on the pipette while ignoring the thermometer is the most common way to waste an evaluation.
Improvements that would actually work
An improvement has to do two things: fix the specific problem you just described, and be possible in the lab you were standing in. Suggesting a bomb calorimeter is not an improvement, it is a wish.
🧩 Turning a weakness into a real improvement
Name the mechanism you are blocking — heat leaving through the top of the cup.
Say what you would change — a lid with a small hole for the thermometer.
Say why that works — it stops evaporation and convection from the surface.
Check it is available — could you do this next lesson with what is in the cupboard?
Weak improvement
Why it fails
A version that works
“Be more careful”
Not a change to the method, and not checkable
Use a pH probe instead of an indicator, so the endpoint is read rather than judged
“Repeat the experiment more times”
Only helps random error, and your problem was systematic
Insulate the calorimeter and fit a lid to cut the heat loss itself
“Use better equipment”
Names nothing, so it fixes nothing
Use a thermometer reading to 0.01 °C, since the temperature rise carried 92% of the uncertainty
“Use a bomb calorimeter”
Not available in a school laboratory
Extrapolate a cooling curve back to the moment of mixing to estimate the true temperature rise
Never write “human error”. It says nothing, and examiners have read it ten thousand times. There is always a specific version underneath it: the endpoint was judged by eye and the pale pink was hard to catch; the stopwatch was started after the reaction had begun; the meniscus was read from above. Write that instead.
Judge the strength of your own support
One last thing, and it is easy to forget. Even when your data supports the hypothesis, say how strongly. Five points sitting neatly on a line with small error bars is strong support. Five points with a visible trend and error bars overlapping each other is weaker support for the same conclusion, and saying so is honest rather than damaging.
WORKED EXAMPLE
Classify each of these as systematic or random, and state whether repeating trials would help: (a) the burette was rinsed with distilled water but not with the acid before filling; (b) the pale pink endpoint was hard to spot and was judged slightly differently each time; (c) the balance had not been zeroed and read 0.02 g high all lesson.
(a) rinsing with waterThe leftover water dilutes the acid slightly, in the same direction every single titration, so every titre comes out a little too large.systematic — repeats will not help(b) judging the endpointSometimes you stop a fraction early, sometimes a fraction late. The results scatter either side of the truth.random — repeats and a mean will help(c) the unzeroed balanceEvery mass is 0.02 g too high by exactly the same amount, all day.systematic — only zeroing the balance helpsTwo out of three cannot be repeated away, which is fairly typical.
WORKED EXAMPLE
Student B measured the enthalpy of neutralisation as −52.3 ± 2 kJ mol−1 against a literature value of −57.3 kJ mol−1, an error of 8.7% against an uncertainty of 3.2%. Write one full evaluation point.
WeaknessThe polystyrene cup was open to the air and stood directly on the bench, so heat escaped from the mixture during the reaction and while the temperature was being read.Impact, with a directionHeat leaving the mixture makes the measured temperature rise smaller than the true one. Since q = mcΔT, a smaller ΔT gives a smaller q and therefore a ΔH that is less exothermic than it should be.every value too small in the same direction: 8.7% error against 3.2% uncertaintythe size and direction both fit heat lossImprovementFit a lid with a hole just large enough for the thermometer, and stand the cup inside a second polystyrene cup. Both cut the escape routes rather than trying to average the problem away.Better stillRecord the temperature every 30 s from before mixing until well after the peak, then extrapolate the cooling section back to the moment of mixing. That estimates the temperature rise there would have been with no heat loss at all.
WORKED EXAMPLE
A student’s uncertainty came out as 2.94% from the temperature rise, 0.12% from the mass and 0.12% from the pipetted volume. Their evaluation opens with a paragraph about the pipette. Explain what is wrong with that choice and what they should write instead.
What is wrongThe pipette contributes 0.12 of a 3.18% total. Even a perfect pipette would barely move the final answer, so a paragraph about it changes nothing.2.94 ÷ 3.18 = 92% of the uncertainty came from one measurementWhat to write insteadLead on the temperature rise. It was only 6.8 °C, and a thermometer reading to ±0.1 °C over two readings gives ±0.2 °C on a small number.a small ΔT is the real weaknessTwo improvements that followUse a thermometer or temperature probe with a finer resolution, or make the temperature rise bigger by using more concentrated solutions — the same absolute uncertainty over a larger ΔT is a smaller percentage.This is why the processing page was worth doing carefully. The uncertainty budget hands you your evaluation, ranked.
💡 Exam tip
Give two or three developed points rather than a long list of small ones.
Use the structure weakness, impact, improvement every time, in that order.
Always state the direction: did the error make your value too high or too low?
Say whether each error is systematic or random, and let that decide the fix.
Include at least one limitation and one assumption — they are separate marks.
Refer to your own numbers: your error bars, your percentage uncertainty, your observations.
⚠️ Common mix-up
Writing “human error” instead of naming what a human actually did.
Suggesting more repeats for a problem that is clearly systematic.
Naming a weakness and jumping straight to the fix, skipping the impact.
Saying “this affected my results” without saying in which direction.
Evaluating the smallest source of uncertainty because it is the easiest to describe.
Writing a generic evaluation that would fit any experiment in the building.
Up next: Exploring — and that is not a mistake. The best research questions come out of an evaluation, because you now know exactly which variable was hard to control and what you would want to test next. The cycle starts again.
Want this explained one-to-one?
Book a free session with an experienced IB Chemistry tutor and get your trickiest topics made simple.