IB Psychology SL Topic 5 — Research Design Paper 1 & 2 Core idea ~10 min read

Reliability: Getting the Same Result Twice

Reliability is consistency, and nothing more than that. If you ran the study again, in the same way, would you get the same sort of result? A reliable measure does not wobble. Notice what that does not promise: a measure can be perfectly consistent and consistently wrong.

📘 What you need to know

What reliability actually means

Think of a set of bathroom scales. Stand on them five times in a minute and they read 62 kg every time. Those scales are reliable. Whether they are right is a completely different question — they might be four kilograms out and still perfectly consistent.

Psychology works the same way. A questionnaire that gives the same person the same anxiety score in January and June is reliable. Whether it measures anxiety at all is a validity question, and it comes next.

The one-line definition Reliability = consistency.
Validity = accuracy.

Why a standardised procedure matters so much

If two researchers run the same study but one reads the instructions warmly and slowly while the other rattles through them, the two runs are not really the same study. Standardisation means every participant meets the same words, the same materials, the same room and the same timing. That is what makes it possible to repeat the study and check the finding.

Which is why lab experiments win here. Controlled conditions, a fixed script, random allocation and quantitative data all point in the same direction: the study can be run again and the numbers compared. Field experiments are harder to repeat, and natural experiments cannot be repeated at all, because nobody controls the event.

Internal and external reliability

These two words trip people up constantly, so hold onto the difference in one sentence: internal is about the parts of the measure agreeing with each other, external is about the measure agreeing with itself later.

Two ways to check a questionnaire is consistent one tests it across time, the other tests it against itself TEST-RETEST external reliability 6 months January June do the scores match? same people, same questions SPLIT-HALF internal reliability odd items even items one questionnaire, cut in two do the halves agree? one sitting, no waiting a strong positive correlation is the evidence in both cases weak agreement means the measure itself needs rewriting
Test-retest needs a gap long enough that people have forgotten their answers, but short enough that the thing being measured has not genuinely changed.
Naming the check is worth a mark on its own. If a question mentions the same test given twice, say test-retest and say external reliability. If it mentions comparing two halves of one questionnaire, say split-half and internal reliability.

Inter-observer reliability

For observations there is a third check. Two trained observers agree the behavioural categories in advance, watch the same session, record separately so neither drifts towards the other, and then compare their tallies. A strong positive correlation between the two records means the categories were clear and the recording was consistent.

This matters for more than tidiness. If one person watches alone, there is nothing stopping them recording what they hoped to see. Two independent records make researcher bias much harder to hide.

Reliable is not the same as valid

Consistent, correct, or neither the centre of the target is the thing you were trying to measure reliable, not valid same answer every time, but the wrong answer reliable and valid consistent and correct what you are aiming for neither all over the place unreliable and invalid reliability is necessary for validity, but nowhere near enough there is no fourth target: valid but unreliable cannot happen
The missing fourth case is the interesting one. If a measure gives you a different answer every time, it cannot be hitting the truth, so it cannot be valid.
CheckWhat it testsHow it is doneUsed for
Test-retestExternal reliabilitySame people, same measure, months apartQuestionnaires, psychometric tests
Split-halfInternal reliabilityCompare two halves of the same measureQuestionnaires with many items
Inter-observerConsistency between peopleTwo observers, same session, separate recordsObservations
ReplicationReliability of the whole studyRun the standardised procedure againExperiments

How to improve reliability

🧩 Five fixes you can offer in any exam answer

  1. Standardise the procedure. Written instructions, read the same way, same materials, same order.
  2. Define your categories precisely so two people would tick the same box.
  3. Train the observers or interviewers before data collection starts.
  4. Pilot the measure and remove any item people read differently.
  5. Use more items. A ten-item scale is less affected by one odd response than a two-item one.

Worked examples

WORKED EXAMPLE

Identify the type of reliability being tested

A researcher gives a 30-item stress questionnaire to 80 nurses. She then compares each nurse’s score on the odd-numbered items with their score on the even-numbered items. Identify what she is testing and explain why. [3]

Step 1: What is being compared? Two halves of the same questionnaire, taken at the same time. Step 2: Name the method The split-half method. Step 3: Say which reliability that is It tests whether the items agree with each other, which is consistency within the measure. Internal reliability a strong positive correlation between the halves is the evidence she is looking for
WORKED EXAMPLE

Explain a reliability problem and fix it

A field study of helping behaviour is run by four researchers on four different days. Each explains the task to passers-by in their own words. Explain one reliability problem and suggest how to solve it. [4]

Step 1: Find the inconsistency Four different explanations means four slightly different studies. Step 2: Name the consequence The procedure is not standardised, so differences in the results may come from the researcher rather than the situation. Step 3: Fix it Write one script, train all four researchers to deliver it identically, and pilot it first. Standardise the procedure you could add: extraneous variables such as weather across four days also reduce reliability

💡 Exam tip

⚠️ Common mix-up

Up next: Validity: Measuring What You Meant To — the other half of the pair, and the harder one to get right.

Want this explained one-to-one?

Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.

Book a Free Session →