IB Psychology HLTopic 5 — Research MethodsPaper 3 & IACore skill~10 min read
Observations: Watching Behaviour Directly
An observation is the only method where the researcher changes nothing and simply watches. That sounds easy. It is not — because the moment you write down what you think someone is feeling, you have stopped observing and started guessing.
📚 What you need to know
An observation is non-experimental: nothing is manipulated, so it cannot show cause and effect.
You can only record what is visible. Motives, thoughts and feelings are off limits.
Naturalistic means a real setting with no interference. Controlled means the researcher sets up the situation.
Covert means participants do not know. Overt means they do.
Participant means the researcher joins the group. Non-participant means they stay outside it.
Those last two pairs combine, so any observation is really a mix of both labels.
Behavioural categories are agreed in advance, and two observers compare records to check inter-observer reliability.
The one rule: record what you can see
Imagine you are watching a child in a playground. You may write down “pushed another child”. You may not write down “was angry”, “wanted attention” or “is naturally aggressive”. You did not see any of those things — you inferred them, and an inference is not data.
The observer’s line
You can record the behaviour.
You cannot record the reason for it.
This is why observations always come with behavioural categories agreed before the study starts. A category has to be something two different people would score the same way. “Aggression” is useless. “Hits another child with an open hand” is usable. That sharpening-up process is called operationalisation, and it is the same skill you need for your IA.
Too vague to score
Operationalised category
Why the second one works
Being friendly
Smiles at a newcomer within 10 seconds of them arriving
There is a clear action and a clear time window
Paying attention
Eyes directed at the teacher, counted every 30 seconds
Everyone knows exactly when to tick the box
Being anxious
Touches face, taps foot, or fidgets with an object
Visible actions replace an invisible feeling
Helping
Moves towards the person who dropped something and picks an item up
Two observers would agree on whether it happened
If you cannot imagine ticking a box the exact moment a behaviour happens, your category is not finished yet. Keep cutting until it is a single visible action.
Naturalistic or controlled?
This pair is about the setting. In a naturalistic observation the world is left alone and you watch whatever happens. In a controlled observation the researcher builds the situation, decides the phases, and sometimes brings in an IV as well — Ainsworth’s Strange Situation is the classic example, with its fixed sequence of separations and reunions.
Bandura used a controlled observation and Festinger used a naturalistic one. Neither was wrong. They wanted different things.
Covert or overt, participant or non-participant
These are two separate questions, so they combine into four possible observations. Getting this grid straight is worth easy marks, because exam questions usually give you a study and ask you to label it.
Most real studies sit in one box but borrow arguments from the others. Say which box you are in before you evaluate.
The ethical get-out for covert work is public behaviour. If people would be doing it in a mall, a stadium or an office corridor anyway, watching them is generally acceptable. Covertly watching people in private is not.
Making the recording trustworthy
One observer working alone is one person’s opinion. Two trained observers who agree are evidence. That agreement is called inter-observer reliability, and it is measured by correlating the two observers’ tallies. A strong positive correlation means the categories are clear and the recording is not just personal bias.
🧩 Running an observation properly
Write the categories first. Every one must be a visible action, not a state of mind.
Train the observers together. Watch a practice clip and check you both tick the same boxes.
Record independently. If you can see each other’s sheets you will drift into agreeing out of politeness, and the reliability figure becomes meaningless.
Choose a sampling method. Event sampling counts every time a behaviour occurs. Time sampling records what is happening at fixed intervals.
Compare the tallies. Correlate the two records. A strong positive result means good inter-observer reliability.
Worked examples
WORKED EXAMPLE
Classify the observation and give one ethical concern
A researcher sits in a busy railway station cafe for three afternoons, noting how many people give up their seat to an older passenger. Nobody is told the study is happening. Classify the observation fully and identify one ethical issue.
Step 1: Setting
Nothing was set up, so it is naturalistic.
Step 2: Do participants know?
No, so it is covert.
Step 3: Is the researcher inside the group?
They are sitting apart and not interacting, so it is non-participant.
Step 4: Ethics
No informed consent and no right to withdraw. This is partly defensible because the cafe is a public place and the behaviour would have happened anyway.
Naturalistic, covert, non-participant; lack of informed consentgive all three labels — a “classify” question usually wants every one
WORKED EXAMPLE
Fix the behavioural categories
A student plans to observe “cooperation” in a group task. Their categories are: being helpful, working well together, and good teamwork. Explain why these will not produce reliable data, and rewrite them.
Step 1: Spot the problem
All three describe a judgement, not an action. Two observers would score the same moment differently.
Step 2: Say what that costs
Low inter-observer reliability, so the data cannot be trusted or replicated.
Step 3: Rewrite as visible actions
Passes a resource to another member. Asks a direct question. Waits for someone to finish speaking before talking.
Step 4: Add the counting rule
Event sampling: tally each time one of the three actions occurs during a 10 minute task.
Operationalised categories plus a counting rule“could two strangers score this identically?” is the test to apply
💡 Exam tip
Give the full label when classifying: naturalistic or controlled, covert or overt, participant or non-participant. Three words, three chances at credit.
Link every strength to a validity term. Covert goes with ecological validity; controlled goes with reliability and replicability.
If a question mentions two observers, it wants inter-observer reliability. Say how it is checked, not just that it exists.
Observations describe. They never explain. Never write that an observation showed one thing “caused” another.
For your IA, operationalised categories are the difference between a top band and a middle band method section.
Name a real study. Festinger, Rosenhan, Bandura and Ainsworth each illustrate a different box on the grid.
⚠ Common mix-up
Treating covert and participant as the same thing. They are separate questions. You can join a group openly, and you can hide while staying outside it.
Calling a field experiment an observation. If an IV is being manipulated, it is an experiment, even if the researcher is watching.
Recording feelings. “Looked nervous” is an inference. “Tapped foot” is an observation.
Assuming naturalistic always means covert. It usually is, but participants can know they are being watched in a real setting.
Saying inter-observer reliability improves validity. It improves reliability. Those are different words and examiners notice.
Claiming covert observation is automatically unethical. In genuinely public settings it is often the accepted approach.
Up next: Questionnaires and Survey Design — what happens when you stop watching people and start asking them instead.
Want this explained one-to-one?
Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.