IB Psychology SLTopic 5 — Methods of ResearchPaper 1 & 2Core skill~10 min read
Observations: Watching Behaviour Directly
An observation is what you use when asking people what they do would ruin the answer. Nobody tells you honestly how often they interrupt in a meeting, but you can sit there and count it. The catch is that you can only ever record what you see — never why.
📘 What you need to know
An observation records observable behaviour. It cannot record motive, feeling or thought.
Naturalistic = a real setting, nothing manipulated. Controlled = the researcher sets up the situation.
Covert = participants do not know. Overt = they do.
Participant = the researcher joins the group. Non-participant = the researcher stays outside it.
Behavioural categories must be agreed and defined before you start watching.
Inter-observer reliability is how closely two trained observers’ records agree.
Covert work buys natural behaviour but costs consent, so it raises ethical problems every time.
What an observation can and cannot say
Imagine you watch a child in a playground hit another child. You can write down: hit, once, with an open hand, after being pushed. What you cannot write down is “because he is an aggressive child” or “because he was angry”. You did not see anger. You saw a hand move.
This sounds obvious and it is where most marks are lost. Observations give you behaviour, and behaviour has to be linked to the theory afterwards, carefully, with no assumption of cause and effect.
The line you cannot cross
You may record what happened. You may not record why it happened.
A good test before you write a category down: could two different people, watching the same ten seconds, both tick it? “Shouts” passes. “Is being rude” does not, because rudeness lives in your head, not on the video.
Three questions that describe any observation
Observations get labelled with two or three words at once, and students find that confusing. It is easier if you treat it as three separate questions. Answer all three and you have named the type.
A study can be covert, naturalistic and participant all at once. Naming all three in an exam answer shows you understand the design rather than a label.
Naturalistic and controlled
A naturalistic observation happens where the behaviour normally happens, with nothing set up: shoppers choosing between two brands, a crowd at a match, children in a playground. There is no IV at all.
You choose it when running an experiment would wreck the very thing you want to see. You cannot randomly allocate children to “have friends” or “be excluded”, so you watch instead.
The strength is the behaviour itself. Nobody is performing, so what you record is unforced. That is high ecological validity, and it is the single best thing about this method.
A controlled observation puts the behaviour in a setting the researcher designed: a room with a one-way mirror, a fixed set of toys, a script for the adult. Because the situation repeats exactly, another researcher can run it again, which is why controlled observations have far better reliability. The price is that the room is not real life, so ecological validity drops.
Covert and overt
Covert
Participants do not know they are being watched.
Behaviour is real and uncontrived
No demand characteristics
No informed consent, no right to withdraw
Only defensible in public places
Overt
Participants know, and were usually told in advance.
Consent and withdrawal are possible
Debriefing is straightforward
Behaviour may be altered by being watched
Opens the door to participant reactivity
The ethical line usually drawn is this: a covert observation is acceptable only where the behaviour would have been visible to any member of the public anyway — a mall, a station, a stadium. Once you are recording behaviour people believed was private, you have a serious problem.
🤔 Why covert work keeps producing famous studies
The moment people know they are being measured, they start managing how they look. Remove that awareness and you get behaviour nobody would have volunteered. That is exactly why covert studies of bystander behaviour and of institutional life have been so influential — and exactly why they are so hard to justify or repeat today.
Participant and non-participant
In a participant observation the researcher joins the group and becomes one of them. This gets you inside conversations, jokes and unspoken rules that an outsider would never see, which is a real gain in validity. Two things then go wrong over time: you only see the part of the group you are standing in, and you start to like the people you are studying. That second one has a name — loss of objectivity — and it is a strong evaluation point.
In a non-participant observation the researcher stays outside, often behind a one-way mirror or at a distance. You keep your objectivity and you usually get a better view of the whole group, but you lose the inside detail, and you cannot ask anyone what just happened.
Behavioural categories: the bit that decides everything
Before a single minute of watching, the researchers agree a list of behaviours and write down exactly what does and does not count. Then they tally them. Vague categories are the reason two observers end up with different numbers.
Category
Counts as this
Does not count
Interrupts
Starts speaking while another person is mid-sentence
Speaking after a pause of more than one second
Off-task
Eyes away from the worksheet for over five seconds
Briefly glancing up and returning
Helps
Moves towards the person and offers words or hands
Looking at the person without moving
Withdraws
Leaves the group area and does not return within a minute
Stepping back but staying in the circle
🧩 Setting up an observation that will actually work
Decide the categories. Small in number, clearly defined, no overlap.
Write the definitions down so both observers read the same words.
Choose a sampling method. Event sampling counts every time it happens; time sampling records what is happening at fixed moments.
Record independently. If the two observers can see each other’s sheets, they will drift towards agreeing.
Compare afterwards and calculate the agreement between them.
Inter-observer reliability
One person watching alone might simply be seeing what they hoped to see. So two trained observers record the same behaviour separately, and the two sets of tallies are compared. A strong positive correlation between them means the categories were clear and the recording was consistent.
Notice the two observers disagree slightly on “withdraws”. That is the category whose definition needs tightening, not the observer who needs replacing.
Worked examples
WORKED EXAMPLE
Name the type of observation
A researcher joins a weekly running club for three months. The other members believe she is simply a new runner. She writes up her notes each evening about how the group encourages slower members. Identify the type of observation and give one limitation of it. [4]
Step 1: Where?
A real running club, nothing set up = naturalisticStep 2: Do they know?
No, they think she is a member = covertStep 3: Is she in the group?
Yes, she runs with them = participantCovert, naturalistic, participant observationlimitation: over three months she may start to identify with the group and lose objectivity, which weakens the validity of her notes
WORKED EXAMPLE
Improve a weak behavioural category
Two observers record “aggression” in a school playground. Observer A records 14 incidents, Observer B records 31. Explain what has gone wrong and how the researchers should fix it. [4]
Name the problem
Poor inter-observer reliability: the same session produced very different counts.
Find the cause“Aggression” is not an observable behaviour — it is a judgement, so each observer used their own threshold.
Fix it
Replace it with defined behaviours: pushes, kicks, takes an object by force, raises voice at another child.
Define behaviours, then re-test agreementthe fix is always in the category definition, not in watching harder
💡 Exam tip
Give the full label. “Covert naturalistic non-participant observation” scores better than “an observation”.
If asked to design an observation, always mention behavioural categories and inter-observer reliability. They are easy marks.
Link every strength and weakness to a named idea: ecological validity, ethical validity, reliability, objectivity.
For ethics, name the specific issue — lack of informed consent, no right to withdraw, no debrief — not just “it is unethical”.
Remember observations are non-experimental. No IV means no cause and effect, ever.
If a study is described as being watched through a one-way mirror, that is non-participant, and usually controlled and overt.
⚠️ Common mix-up
Covert vs participant. They often go together but they answer different questions. A researcher can join a group openly.
Controlled observation vs lab experiment. An observation may have no IV at all; if there is a manipulated IV, call it an experiment.
Writing motives into the record. “Looked bored” is an inference. “Yawned, put head on desk” is an observation.
Confusing reliability with validity. Two observers agreeing perfectly on a badly chosen category is reliable and still meaningless.
Assuming naturalistic means ethical. Naturalistic and covert together is usually where the ethical trouble is.
Thinking a tally chart is the method. The tally chart is just the recording tool.
Up next: Questionnaires and Survey Design — what happens when you stop watching people and start asking them.
Want this explained one-to-one?
Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.