IB Psychology HLTopic 5 — Research DesignPaper 3 & IACore idea~9 min read
Reliability: Getting the Same Result Twice
Reliability is consistency, nothing more. If you weigh yourself three times in a minute and get three different numbers, the scale is unreliable — even if one of those numbers happened to be right. Research works the same way, and this is the idea everything else in research design is built on.
📚 What you need to know
Reliability means a measure gives consistent results when it is used again.
It comes from a standardised procedure: the same instructions, materials and conditions every time.
Internal reliability is consistency within a measure. External reliability is consistency over time.
Inter-observer reliability is agreement between two or more trained observers watching the same thing.
Lab experiments are the most reliable method, because control and standardisation are highest.
A reliable measure is not automatically valid. You can be consistently wrong.
Where reliability comes from
It comes from doing the same thing every single time. That is what a standardised procedure means: identical instructions read to every participant, identical materials, identical room, identical timing. The moment one participant gets a friendlier greeting or an extra thirty seconds, you have introduced a difference that has nothing to do with your study, and repeating it becomes impossible.
The one-line definition
Reliability = would you get the same result again?
This is why lab experiments score highest on reliability. Everything is controlled, the procedure is written down, and another researcher can copy it exactly. Field experiments lose some of this because the real world will not hold still. Natural experiments lose more, because the IV was an event nobody controlled. Unstructured interviews and naturalistic observations sit at the bottom — not because they are bad, but because they were never designed to be repeated.
The two ways to check a measure
The gap in test-retest matters. Too short and people remember their answers; too long and they may genuinely have changed. Six months is the usual compromise.
Both checks end in the same place: a correlation. If you are asked how reliability is measured, “the two sets of scores are correlated, and a strong positive correlation shows good reliability” is the sentence that earns the mark.
Inter-observer reliability
Observations have their own version of this problem. One person’s tally chart is one person’s judgement. So two trained observers watch the same behaviour at the same time, record independently, and their tallies are compared. If they agree closely, the behavioural categories are clear enough for anyone to use, and researcher bias has less room to operate.
Notice that the two observers here do not match perfectly on frowns. Real agreement is never total, which is why it is expressed as a correlation rather than a yes or no.
The trap worth remembering: a broken thermometer that always reads two degrees high is perfectly reliable. It gives the same answer every time. It is also wrong every time. Reliability guarantees consistency, never accuracy.
🧩 Building reliability into your own study
Write a script. Read the same instructions word for word to every participant.
Fix the conditions. Same room, same time of day, same materials, same time limit.
Use standardised measures where possible — an established scale has usually been reliability tested already.
Pilot the procedure and rewrite anything that had to be explained twice.
If you are observing, use two observers and report their agreement.
Record what you did in enough detail that a stranger could repeat it. That is the real test.
Worked examples
WORKED EXAMPLE
Explain how reliability could be assessed
A psychologist develops a 16-item questionnaire measuring exam anxiety. Explain one way she could assess its reliability, and state what result would show the questionnaire is reliable.
Step 1: Choose a suitable method
Use the split-half method to check internal reliability.
Step 2: Say how it is done
Split the 16 items into two sets of 8 — for example odd-numbered and even-numbered items — and calculate a separate score for each half for every participant.
Step 3: Say how it is judgedCorrelate the two sets of half-scores.
Step 4: State the expected result
A strong positive correlation shows all the items are measuring the same construct, so internal reliability is high.
Split-half, then correlate; strong positive correlation = reliable“how it is assessed” wants the procedure and the result, not just a name
WORKED EXAMPLE
Reliable but not valid
A school uses a five-minute online quiz to measure “student wellbeing”. Students who take it twice in a term get almost identical scores. The head of year argues this proves the quiz is a good measure. Evaluate that claim.
Step 1: Name what the evidence actually shows
Matching scores across time show high external reliability (test-retest).
Step 2: Name what it does not show
It says nothing about validity — whether the quiz measures wellbeing at all.
Step 3: Give the mechanism
A quiz that consistently measures the wrong thing, for example tiredness, will still give consistent scores.
Step 4: Say what would be needed
Evidence of construct validity: does the score agree with an established wellbeing measure or with clinical judgement?
Reliable, but reliability alone cannot establish validitythis pairing comes up constantly — consistent is not the same as correct
💡 Exam tip
Learn the pairing: split-half is internal, test-retest is external. Get them the wrong way round and the whole answer collapses.
Always finish with the correlation. “The two sets of scores are correlated” is the mechanism the mark scheme wants.
Standardised procedure is the phrase to use whenever a question asks how reliability could be improved.
If a study uses observers, inter-observer reliability is almost always a creditable point.
Never write that a reliable study must be valid. The reverse claim is safer: a valid study does need to be reliable.
Reliability applies to measures and procedures, not to findings. Keep the wording precise.
⚠ Common mix-up
Using reliability to mean trustworthy. In everyday English it does. In psychology it only means consistent.
Swapping internal and external. Internal = within the measure; external = across time.
Thinking test-retest needs a different test. It is the same test, same people, later.
Confusing inter-observer reliability with validity. Two observers agreeing shows consistency, not accuracy.
Assuming qualitative research is simply unreliable. It is judged by credibility instead, which is a different standard, not a lower one.
Writing “the results were reliable”. Say the measure, procedure or observation was reliable.
Up next: Validity: Measuring What You Meant To — the other half of the pair, and the one that decides whether the study was worth doing at all.
Want this explained one-to-one?
Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.