IB Physics SLInquiry 3 — Concluding & EvaluatingInternal AssessmentRandom vs systematic errors~9 min read
Evaluating the Method
The evaluation is where you turn a critical eye on your own experiment. Not to apologise for it — to show you understand why it behaved the way it did. You’ll pin down specific sources of error, separate the random from the systematic, weigh up your method’s weaknesses and limits, and propose fixes that could actually be done in a school lab. This is often the highest-scoring section of an IA, and the one where “human error” quietly loses marks.
📘 What you need to know
Revisit your hypothesis and judge the strength of the support, given your errors
Identify specific sources of error — never the phrase “human error”
Random errors scatter results (hurt precision); systematic errors shift them one way (hurt accuracy)
Discuss your method’s weaknesses, limitations, and assumptions separately — they mean different things
For every weakness, close the loop: Weakness → Impact → Improvement
Improvements must be relevant (fix the actual issue) and realistic (doable in a school lab)
Evaluate the hypothesis
Start by looping back to your hypothesis, following on from your conclusion. Even when the data supported it, judge how strong that support really is in light of the errors you found. For instance: the data backed the claim that resistance is proportional to length — but a y-intercept that didn’t pass through the origin hints at a systematic error, which slightly weakens the confirmation of the theoretical model. That kind of honest nuance scores well.
Random vs systematic errors
Identifying and discussing errors is the most important part of the evaluation — and the first job is telling the two types apart, because they behave differently and are fixed differently.
Random error scatters points around the true value; systematic error shifts a tight cluster consistently to one side.
Random errors
Random errors are unpredictable variations that happen by chance, scattering results around the true value and harming precision. Classic examples: fluctuating reaction time when starting and stopping a stopwatch; parallax when reading an analogue scale from an angle; a reading error from a ruler marked only in centimetres.
You minimise random error by taking repeat trials and averaging, measuring over many intervals (like timing 20 oscillations and dividing), and using a more precise instrument with finer divisions.
repeat & average
time many swings
finer instrument
→ all cut →
random error
Systematic errors
Systematic errors are flaws in the method or apparatus that push every reading the same way — always too high or always too low — harming accuracy. Examples: an uncorrected zero error on a balance or ammeter that offsets every reading; not accounting for heat loss in a thermal experiment, which always makes the measured temperature change too small.
You reduce systematic error by recalibrating the apparatus, switching to different apparatus, or correcting the method itself. On a graph, a systematic error shows up as a line that’s offset from the origin — parallel to where it should be, but shifted.
A systematic error keeps the gradient but lifts the whole line off the origin — the constant offset is the tell.
Weaknesses, limitations, and assumptions
Beyond errors, three related ideas deserve separate treatment — they’re often muddled together, but each means something distinct.
🧭 Three things to discuss — kept apart
Weaknesses — aspects of the method that cause significant errors. E.g. timing a single swing of a pendulum, where the short period makes reaction-time error huge.
Limitations — factors that bound how far your conclusion applies. E.g. only testing up to 1.00 m, so you can’t claim T2 ∝ L holds for very long pendulums.
Assumptions — simplifications in your calculations that aren’t perfectly true. E.g. assuming air resistance is negligible and the string is massless.
Close the loop: Weakness → Impact → Improvement
This is the structure examiners look for. For every weakness you name, explain its impact on your final result, then propose a specific improvement that fixes it. An improvement must be relevant (it addresses the actual weakness) and realistic (you could do it in a normal school lab — a bomb calorimeter doesn’t count).
Weakness
→ causes →
Impact on result
→ fixed by →
Improvement
Quick recap: evaluate the hypothesis honestly, separate random (scatter) from systematic (shift) errors, keep weaknesses / limitations / assumptions distinct, and close every loop with a relevant, realistic fix.
WE 1
In a pendulum experiment to find g, the length was measured to the bottom of the bob, not its centre of mass. Evaluate this as Weakness → Impact → Improvement.
Weakness (systematic)
length measured to the bob’s bottom, not its centre
Impact
measured L is consistently too long
since T ∝ √L, every period comes out a bit large
→ calculated g is systematically too highImprovement (relevant + realistic)
measure to the geometric centre of the bob
add half the bob’s diameter (vernier callipers) to the string length
WE 2
In the same experiment, the period was found by timing 20 swings with a manual stopwatch. Evaluate this weakness and its fix.
Weakness (random)
human reaction time starting/stopping the stopwatch
Impact
adds scatter to the measured times
visible as error bars and points scattered around the line of best fit — a precision problem, not an accuracy one.Improvement
use a light gate at the bottom of the swing
automatically times the period, removing reaction-time error
💡 Top tips
Be specific — never “human error”. Name the mechanism: “parallax when reading the pointer position could have skewed the length measurements”.
Prioritise the one or two errors with the biggest impact on your result — for a pendulum, the period and effective length matter far more than pivot friction.
Always close the loop — Weakness → Impact → Improvement, every time.
Evaluate your own data — refer to your actual graph and observations (e.g. “the non-zero intercept on the R vs. L graph”), not a generic template.
⚠ Common mistakes
Blaming vague “human error” instead of a named, specific source
Confusing random (scatter, precision) with systematic (shift, accuracy)
Listing a weakness but not stating its impact or suggesting a fix
Proposing an unrealistic improvement you couldn’t do in a school lab
That completes Inquiry 3: Concluding & Evaluating — and with it, the whole Scientific Inquiry Cycle. You can now run an investigation end to end: explore and design it, collect and process the data, interpret and conclude, then evaluate it with a clear-eyed, specific critique. That’s exactly the arc a top-band IA follows.
Want this to actually click before the exam?
Book a free meeting and let’s work through the tricky bits together.