IB Psychology SLTopic 5 — Research DesignPaper 1 & 2Core idea~10 min read
Reliability: Getting the Same Result Twice
Reliability is consistency, and nothing more than that. If you ran the study again, in the same way, would you get the same sort of result? A reliable measure does not wobble. Notice what that does not promise: a measure can be perfectly consistent and consistently wrong.
📘 What you need to know
Reliability means consistency: the same procedure gives similar results when repeated.
A standardised procedure is what makes a study repeatable in the first place.
Internal reliability: the measure is consistent within itself.
External reliability: the measure is consistent over time.
Inter-observer reliability checks that two trained observers record the same thing.
Lab experiments are the most reliable method; natural experiments the least, because nothing was controlled.
What reliability actually means
Think of a set of bathroom scales. Stand on them five times in a minute and they read 62 kg every time. Those scales are reliable. Whether they are right is a completely different question — they might be four kilograms out and still perfectly consistent.
Psychology works the same way. A questionnaire that gives the same person the same anxiety score in January and June is reliable. Whether it measures anxiety at all is a validity question, and it comes next.
The one-line definition
Reliability = consistency. Validity = accuracy.
Why a standardised procedure matters so much
If two researchers run the same study but one reads the instructions warmly and slowly while the other rattles through them, the two runs are not really the same study. Standardisation means every participant meets the same words, the same materials, the same room and the same timing. That is what makes it possible to repeat the study and check the finding.
Which is why lab experiments win here. Controlled conditions, a fixed script, random allocation and quantitative data all point in the same direction: the study can be run again and the numbers compared. Field experiments are harder to repeat, and natural experiments cannot be repeated at all, because nobody controls the event.
Internal and external reliability
These two words trip people up constantly, so hold onto the difference in one sentence: internal is about the parts of the measure agreeing with each other, external is about the measure agreeing with itself later.
Test-retest needs a gap long enough that people have forgotten their answers, but short enough that the thing being measured has not genuinely changed.
Naming the check is worth a mark on its own. If a question mentions the same test given twice, say test-retest and say external reliability. If it mentions comparing two halves of one questionnaire, say split-half and internal reliability.
Inter-observer reliability
For observations there is a third check. Two trained observers agree the behavioural categories in advance, watch the same session, record separately so neither drifts towards the other, and then compare their tallies. A strong positive correlation between the two records means the categories were clear and the recording was consistent.
This matters for more than tidiness. If one person watches alone, there is nothing stopping them recording what they hoped to see. Two independent records make researcher bias much harder to hide.
Reliable is not the same as valid
The missing fourth case is the interesting one. If a measure gives you a different answer every time, it cannot be hitting the truth, so it cannot be valid.
Check
What it tests
How it is done
Used for
Test-retest
External reliability
Same people, same measure, months apart
Questionnaires, psychometric tests
Split-half
Internal reliability
Compare two halves of the same measure
Questionnaires with many items
Inter-observer
Consistency between people
Two observers, same session, separate records
Observations
Replication
Reliability of the whole study
Run the standardised procedure again
Experiments
How to improve reliability
🧩 Five fixes you can offer in any exam answer
Standardise the procedure. Written instructions, read the same way, same materials, same order.
Define your categories precisely so two people would tick the same box.
Train the observers or interviewers before data collection starts.
Pilot the measure and remove any item people read differently.
Use more items. A ten-item scale is less affected by one odd response than a two-item one.
Worked examples
WORKED EXAMPLE
Identify the type of reliability being tested
A researcher gives a 30-item stress questionnaire to 80 nurses. She then compares each nurse’s score on the odd-numbered items with their score on the even-numbered items. Identify what she is testing and explain why. [3]
Step 1: What is being compared?
Two halves of the same questionnaire, taken at the same time.
Step 2: Name the methodThe split-half method.Step 3: Say which reliability that is
It tests whether the items agree with each other, which is consistency within the measure.
Internal reliabilitya strong positive correlation between the halves is the evidence she is looking for
WORKED EXAMPLE
Explain a reliability problem and fix it
A field study of helping behaviour is run by four researchers on four different days. Each explains the task to passers-by in their own words. Explain one reliability problem and suggest how to solve it. [4]
Step 1: Find the inconsistency
Four different explanations means four slightly different studies.
Step 2: Name the consequenceThe procedure is not standardised, so differences in the results may come from the researcher rather than the situation.
Step 3: Fix it
Write one script, train all four researchers to deliver it identically, and pilot it first.
Standardise the procedureyou could add: extraneous variables such as weather across four days also reduce reliability
💡 Exam tip
Define reliability as consistency in your first sentence. It frames everything after it.
Always say which reliability: internal, external, or inter-observer.
Standardisation is the answer to most “how could reliability be improved” questions.
Link method to reliability directly: lab high, field lower, natural lowest, and say why.
Qualitative research does not usually claim reliability. It claims credibility instead, which comes later in this topic.
If you mention test-retest, mention the gap. Too short and people just remember their answers.
⚠️ Common mix-up
Using reliable to mean trustworthy. In psychology it means consistent, nothing else.
Swapping internal and external. Internal = within the measure. External = across time.
Thinking reliable means valid. Consistently wrong is still consistent.
Confusing replication with repetition. Replication means another researcher runs it again, which is the stronger test.
Saying a study is unreliable because the sample was small. That is a generalisability problem, not a reliability one.
Forgetting inter-observer reliability exists. For observations it is the main check.
Up next: Validity: Measuring What You Meant To — the other half of the pair, and the harder one to get right.
Want this explained one-to-one?
Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.