IB Psychology SL Topic 5 — Methods of Research Paper 1 & 2 Core skill ~10 min read

Observations: Watching Behaviour Directly

An observation is what you use when asking people what they do would ruin the answer. Nobody tells you honestly how often they interrupt in a meeting, but you can sit there and count it. The catch is that you can only ever record what you see — never why.

📘 What you need to know

What an observation can and cannot say

Imagine you watch a child in a playground hit another child. You can write down: hit, once, with an open hand, after being pushed. What you cannot write down is “because he is an aggressive child” or “because he was angry”. You did not see anger. You saw a hand move.

This sounds obvious and it is where most marks are lost. Observations give you behaviour, and behaviour has to be linked to the theory afterwards, carefully, with no assumption of cause and effect.

The line you cannot cross You may record what happened.
You may not record why it happened.
A good test before you write a category down: could two different people, watching the same ten seconds, both tick it? “Shouts” passes. “Is being rude” does not, because rudeness lives in your head, not on the video.

Three questions that describe any observation

Observations get labelled with two or three words at once, and students find that confusing. It is easier if you treat it as three separate questions. Answer all three and you have named the type.

Three questions that describe any observation answer all three and you have named the type WHERE DOES IT HAPPEN? Naturalistic Controlled DO THEY KNOW? Covert Overt IS THE RESEARCHER IN THE GROUP? Participant Non-participant most studies sit at one end of each scale, so you get a three-word label
A study can be covert, naturalistic and participant all at once. Naming all three in an exam answer shows you understand the design rather than a label.

Naturalistic and controlled

A naturalistic observation happens where the behaviour normally happens, with nothing set up: shoppers choosing between two brands, a crowd at a match, children in a playground. There is no IV at all.

You choose it when running an experiment would wreck the very thing you want to see. You cannot randomly allocate children to “have friends” or “be excluded”, so you watch instead.

The strength is the behaviour itself. Nobody is performing, so what you record is unforced. That is high ecological validity, and it is the single best thing about this method.

A controlled observation puts the behaviour in a setting the researcher designed: a room with a one-way mirror, a fixed set of toys, a script for the adult. Because the situation repeats exactly, another researcher can run it again, which is why controlled observations have far better reliability. The price is that the room is not real life, so ecological validity drops.

Covert and overt

Covert

Participants do not know they are being watched.

  • Behaviour is real and uncontrived
  • No demand characteristics
  • No informed consent, no right to withdraw
  • Only defensible in public places

Overt

Participants know, and were usually told in advance.

  • Consent and withdrawal are possible
  • Debriefing is straightforward
  • Behaviour may be altered by being watched
  • Opens the door to participant reactivity

The ethical line usually drawn is this: a covert observation is acceptable only where the behaviour would have been visible to any member of the public anyway — a mall, a station, a stadium. Once you are recording behaviour people believed was private, you have a serious problem.

🤔 Why covert work keeps producing famous studies

The moment people know they are being measured, they start managing how they look. Remove that awareness and you get behaviour nobody would have volunteered. That is exactly why covert studies of bystander behaviour and of institutional life have been so influential — and exactly why they are so hard to justify or repeat today.

Participant and non-participant

In a participant observation the researcher joins the group and becomes one of them. This gets you inside conversations, jokes and unspoken rules that an outsider would never see, which is a real gain in validity. Two things then go wrong over time: you only see the part of the group you are standing in, and you start to like the people you are studying. That second one has a name — loss of objectivity — and it is a strong evaluation point.

In a non-participant observation the researcher stays outside, often behind a one-way mirror or at a distance. You keep your objectivity and you usually get a better view of the whole group, but you lose the inside detail, and you cannot ask anyone what just happened.

Behavioural categories: the bit that decides everything

Before a single minute of watching, the researchers agree a list of behaviours and write down exactly what does and does not count. Then they tally them. Vague categories are the reason two observers end up with different numbers.

CategoryCounts as thisDoes not count
InterruptsStarts speaking while another person is mid-sentenceSpeaking after a pause of more than one second
Off-taskEyes away from the worksheet for over five secondsBriefly glancing up and returning
HelpsMoves towards the person and offers words or handsLooking at the person without moving
WithdrawsLeaves the group area and does not return within a minuteStepping back but staying in the circle

🧩 Setting up an observation that will actually work

  1. Decide the categories. Small in number, clearly defined, no overlap.
  2. Write the definitions down so both observers read the same words.
  3. Choose a sampling method. Event sampling counts every time it happens; time sampling records what is happening at fixed moments.
  4. Record independently. If the two observers can see each other’s sheets, they will drift towards agreeing.
  5. Compare afterwards and calculate the agreement between them.

Inter-observer reliability

One person watching alone might simply be seeing what they hoped to see. So two trained observers record the same behaviour separately, and the two sets of tallies are compared. A strong positive correlation between them means the categories were clear and the recording was consistent.

Two observers, one set of behaviour agree the categories first, record separately, compare afterwards Observer A Interrupts 7 Off-task 4 Helps 9 Withdraws 2 Observer B Interrupts 8 Off-task 4 Helps 9 Withdraws 3 compare the two records records match closely = good inter-observer reliability when the tallies match, the categories were clear enough to use when they do not, fix the definitions before blaming the observers
Notice the two observers disagree slightly on “withdraws”. That is the category whose definition needs tightening, not the observer who needs replacing.

Worked examples

WORKED EXAMPLE

Name the type of observation

A researcher joins a weekly running club for three months. The other members believe she is simply a new runner. She writes up her notes each evening about how the group encourages slower members. Identify the type of observation and give one limitation of it. [4]

Step 1: Where? A real running club, nothing set up = naturalistic Step 2: Do they know? No, they think she is a member = covert Step 3: Is she in the group? Yes, she runs with them = participant Covert, naturalistic, participant observation limitation: over three months she may start to identify with the group and lose objectivity, which weakens the validity of her notes
WORKED EXAMPLE

Improve a weak behavioural category

Two observers record “aggression” in a school playground. Observer A records 14 incidents, Observer B records 31. Explain what has gone wrong and how the researchers should fix it. [4]

Name the problem Poor inter-observer reliability: the same session produced very different counts. Find the cause “Aggression” is not an observable behaviour — it is a judgement, so each observer used their own threshold. Fix it Replace it with defined behaviours: pushes, kicks, takes an object by force, raises voice at another child. Define behaviours, then re-test agreement the fix is always in the category definition, not in watching harder

💡 Exam tip

⚠️ Common mix-up

Up next: Questionnaires and Survey Design — what happens when you stop watching people and start asking them.

Want this explained one-to-one?

Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.

Book a Free Session →