IB Psychology HL Topic 5 — Research Methods Paper 3 & IA Core skill ~10 min read

Observations: Watching Behaviour Directly

An observation is the only method where the researcher changes nothing and simply watches. That sounds easy. It is not — because the moment you write down what you think someone is feeling, you have stopped observing and started guessing.

📚 What you need to know

The one rule: record what you can see

Imagine you are watching a child in a playground. You may write down “pushed another child”. You may not write down “was angry”, “wanted attention” or “is naturally aggressive”. You did not see any of those things — you inferred them, and an inference is not data.

The observer’s line You can record the behaviour.
You cannot record the reason for it.

This is why observations always come with behavioural categories agreed before the study starts. A category has to be something two different people would score the same way. “Aggression” is useless. “Hits another child with an open hand” is usable. That sharpening-up process is called operationalisation, and it is the same skill you need for your IA.

Too vague to scoreOperationalised categoryWhy the second one works
Being friendlySmiles at a newcomer within 10 seconds of them arrivingThere is a clear action and a clear time window
Paying attentionEyes directed at the teacher, counted every 30 secondsEveryone knows exactly when to tick the box
Being anxiousTouches face, taps foot, or fidgets with an objectVisible actions replace an invisible feeling
HelpingMoves towards the person who dropped something and picks an item upTwo observers would agree on whether it happened
If you cannot imagine ticking a box the exact moment a behaviour happens, your category is not finished yet. Keep cutting until it is a single visible action.

Naturalistic or controlled?

This pair is about the setting. In a naturalistic observation the world is left alone and you watch whatever happens. In a controlled observation the researcher builds the situation, decides the phases, and sometimes brings in an IV as well — Ainsworth’s Strange Situation is the classic example, with its fixed sequence of separations and reunions.

Two settings, two kinds of evidence The more you tidy the situation, the less it looks like real life. NATURALISTIC OBSERVATION • Real setting, nothing set up • People often do not know • High ecological validity • Very hard to replicate • Consent is a real problem Real behaviour, weak control CONTROLLED OBSERVATION • Researcher sets the scene • Fixed phases and categories • Replicable, so more reliable • The setting feels artificial • Demand characteristics likely Tidy control, less realism You cannot have both at full strength in the same study. Pick the one that suits your research question, then admit the cost in your evaluation.
Bandura used a controlled observation and Festinger used a naturalistic one. Neither was wrong. They wanted different things.

Covert or overt, participant or non-participant

These are two separate questions, so they combine into four possible observations. Getting this grid straight is worth easy marks, because exam questions usually give you a study and ask you to label it.

Four ways to watch Do they know you are there? And are you inside the group or outside it? COVERT they do not know they are watched OVERT they know and have agreed PARTICIPANT joins the group NON-PARTICIPANT stays outside Inside, hidden Rosenhan: researchers got themselves admitted as patients − deception, no consent Inside, known A researcher joins a team and everyone knows the role − the group may perform Outside, hidden One-way mirror, or watching a crowd in a public place + behaviour stays natural Outside, known Bandura watched the children from an adjoining room + consent and debrief work Covert buys you honest behaviour and costs you consent. Overt buys you clean ethics and costs you natural behaviour. Say which cost you accepted.
Most real studies sit in one box but borrow arguments from the others. Say which box you are in before you evaluate.
The ethical get-out for covert work is public behaviour. If people would be doing it in a mall, a stadium or an office corridor anyway, watching them is generally acceptable. Covertly watching people in private is not.

Making the recording trustworthy

One observer working alone is one person’s opinion. Two trained observers who agree are evidence. That agreement is called inter-observer reliability, and it is measured by correlating the two observers’ tallies. A strong positive correlation means the categories are clear and the recording is not just personal bias.

🧩 Running an observation properly

  1. Write the categories first. Every one must be a visible action, not a state of mind.
  2. Train the observers together. Watch a practice clip and check you both tick the same boxes.
  3. Record independently. If you can see each other’s sheets you will drift into agreeing out of politeness, and the reliability figure becomes meaningless.
  4. Choose a sampling method. Event sampling counts every time a behaviour occurs. Time sampling records what is happening at fixed intervals.
  5. Compare the tallies. Correlate the two records. A strong positive result means good inter-observer reliability.

Worked examples

WORKED EXAMPLE

Classify the observation and give one ethical concern

A researcher sits in a busy railway station cafe for three afternoons, noting how many people give up their seat to an older passenger. Nobody is told the study is happening. Classify the observation fully and identify one ethical issue.

Step 1: Setting Nothing was set up, so it is naturalistic. Step 2: Do participants know? No, so it is covert. Step 3: Is the researcher inside the group? They are sitting apart and not interacting, so it is non-participant. Step 4: Ethics No informed consent and no right to withdraw. This is partly defensible because the cafe is a public place and the behaviour would have happened anyway. Naturalistic, covert, non-participant; lack of informed consent give all three labels — a “classify” question usually wants every one
WORKED EXAMPLE

Fix the behavioural categories

A student plans to observe “cooperation” in a group task. Their categories are: being helpful, working well together, and good teamwork. Explain why these will not produce reliable data, and rewrite them.

Step 1: Spot the problem All three describe a judgement, not an action. Two observers would score the same moment differently. Step 2: Say what that costs Low inter-observer reliability, so the data cannot be trusted or replicated. Step 3: Rewrite as visible actions Passes a resource to another member. Asks a direct question. Waits for someone to finish speaking before talking. Step 4: Add the counting rule Event sampling: tally each time one of the three actions occurs during a 10 minute task. Operationalised categories plus a counting rule “could two strangers score this identically?” is the test to apply

💡 Exam tip

⚠ Common mix-up

Up next: Questionnaires and Survey Design — what happens when you stop watching people and start asking them instead.

Want this explained one-to-one?

Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.

Book a Free Session →