IB Psychology HLTopic 5 — Research MethodsPaper 3 & IACore skill~11 min read
Experiments and Their Designs
Out of every method in psychology, only the experiment can tell you that one thing caused another. Everything else can spot a link. That is why examiners love it, and why they will happily take marks off you for not knowing the difference between a lab experiment and a quasi-experiment.
📚 What you need to know
In an experiment the researcher changes one variable (the IV) and measures another (the DV), holding everything else steady.
Lab experiments have the tightest control, so they have high internal validity but often feel artificial.
Field experiments happen in a real setting: more realistic, but harder to control.
Natural and quasi-experiments use an IV the researcher cannot manipulate (a war, an illness, someone’s age).
There are three designs: independent measures, repeated measures and matched pairs.
Each design has one signature weakness: participant variables, order effects, or the sheer difficulty of matching people up.
The fixes are random allocation and counterbalancing. Know which one fixes which problem.
What an experiment actually does
Think of a mixing desk with twenty sliders on it. An experiment means putting your hand on one slider, moving it, and watching what happens to the sound. If you moved five sliders at once you would hear a change, but you would have no idea which slider caused it. That is the whole logic of experimental method in one picture.
The three variablesIV = the thing you change • DV = the thing you measure Controlled variables = everything you deliberately keep the same
The red box is the reason experiments feel so fussy. Every control the researcher adds is one more rival explanation shut out.
A variable is only “extraneous” until it actually affects your results. The moment it does, and it lines up with your conditions, it gets a new name: a confounding variable. Use that word in an exam and you sound like you know what you are doing.
The four kinds of experiment
They differ on two things only: how much the researcher controls, and whether the researcher can actually manipulate the IV. Everything else follows from those two.
Type
Setting
Does the researcher manipulate the IV?
Strength / weakness
Laboratory
Controlled space
Yes, fully
High internal validity; task can feel artificial
Field
Real-world setting
Yes, but with less control
High ecological validity; extraneous variables everywhere
Natural
Wherever the event happened
No, the event happened anyway
Lets you study things it would be unethical to cause; no random allocation
Quasi
Often a lab
No, the IV is a feature of the person
Runs like a real experiment; participant variables are baked in
The line students trip over is natural versus quasi. A natural experiment has an IV that is an event nobody could ethically arrange: living through a war, surviving a plane crash, a school being closed for a year. A quasi-experiment has an IV that is a characteristic of the person: their age, their gender, whether or not they already have a diagnosis. In both cases you cannot randomly allocate anybody, and that is exactly why neither can prove cause and effect as cleanly as a lab experiment.
Field experiment or naturalistic observation? If somebody is still changing an IV on purpose, it is a field experiment. Piliavin’s subway study staged a collapse, so the “victim” being drunk or disabled was a deliberately manipulated IV — a field experiment, not an observation.
The three experimental designs
Once you have decided what the IV is, you still have to decide who does what. That is the design, and there are only three options.
Notice the trade. The moment you stop using the same people, participant variables walk in. The moment you reuse them, order effects walk in.
The two problems, and the two fixes
Participant variables are the differences between people that have nothing to do with your study — memory, motivation, how much sleep they had. In an independent measures design, if all the sharper participants happen to land in condition 1, condition 1 will win and you will wrongly credit the IV. The fix is random allocation: decide who goes where by chance, so those differences spread out evenly across both groups.
Order effects are what happens when doing the first condition changes how somebody performs in the second. They get better through practice, or worse through boredom and tiredness. The fix is counterbalancing: half the participants do condition 1 first, the other half do condition 2 first. The practice effect is still there, but it now pushes both conditions equally, so it cancels out.
🧩 Choosing your design in the IA
Can the same person do both conditions without it being obvious? If a participant would clearly work out the aim second time round, go independent measures.
Is the task one you can only really do once? Solving a riddle, hearing a story for the first time. Then it has to be independent measures.
Are participants scarce? Repeated measures gets twice the data from the same people, so it needs fewer of them.
Do you have the time to pre-test everyone? Only then is matched pairs realistic. In a school IA, usually you do not.
Write down the fix. Independent measures needs random allocation. Repeated measures needs counterbalancing. Say so in your method.
Counterbalancing has a nickname: the ABBA design. Group one does A then B, group two does B then A. It does not remove the practice effect — it just shares it out fairly.
Worked examples
WORKED EXAMPLE
Identify the design and its main weakness
A researcher tests whether background music affects reading comprehension. Thirty students read a passage in silence and answer ten questions. The next day the same thirty read a different passage with music playing and answer ten more questions. Identify the IV, the DV and the design, and state one problem with it.
Step 1: Find what was deliberately changed
IV = whether there is background music (silence or music).
Step 2: Find what was measured
DV = score out of 10 on the comprehension questions.
Step 3: Ask who did what
The same thirty students did both conditions, so this is repeated measures.
Step 4: Name the matching weakness
Repeated measures brings order effects — on day two they already know the question style, so they may score higher through practice, not because of the music.
Repeated measures; order effects; fix with counterbalancingsay the fix as well as the problem — that is usually where the second mark sits
WORKED EXAMPLE
Lab, field, natural or quasi?
Researchers compare the working memory scores of children who were living in a flooded region during a major flood with children from a nearby town that was not flooded. All children are tested in the same quiet school room using the same digit span task. Classify this study and explain your choice.
Step 1: Was the IV manipulated by the researcher?
No. Nobody arranged the flood. The IV is flood exposure: yes or no.
Step 2: Is the IV an event or a personal characteristic?
An event that happened anyway, so this is a natural experiment, not a quasi-experiment.
Step 3: Do not be fooled by the setting
Testing in a quiet school room makes it controlled, but control is not what defines the type — manipulation of the IV is.
Step 4: Give the consequence
Children could not be randomly allocated, so differences in income, schooling or family support could explain the result instead.
Natural experiment — unmanipulated IV, no random allocation“they were tested in a lab” does not make it a lab experiment
💡 Exam tip
Write your IV and DV as a full sentence, not one word. “The IV was whether participants heard music or silence while reading” earns credit; “music” does not.
When asked to evaluate a method, always link it to a named validity: internal, ecological, construct. That link is what turns a comment into a mark.
Match the problem to the design. Order effects only exist in repeated measures. Participant variables are the independent measures problem.
Random allocation is not the same as random sampling. Allocation is about which condition; sampling is about who gets in the study at all.
For the HL IA, state your design and your control measures explicitly. Markers look for the words, not just the idea.
If a question says “suggest one improvement”, give a change plus what it would fix. Two halves, two marks.
⚠ Common mix-up
Calling a quasi-experiment a natural experiment. Age, gender and diagnosis are characteristics of the person, so those are quasi. A flood, a war or a hospital closure is an event, so that is natural.
Thinking a lab experiment must involve equipment. A “lab” just means a controlled setting the researcher set up. A quiet classroom counts.
Confusing extraneous with confounding. Extraneous is any nuisance variable. Confounding means it actually varied with the IV and messed up your result.
Saying counterbalancing “removes” order effects. It balances them across conditions. It does not delete them.
Writing that field experiments have no control. They have less control. The IV is still manipulated.
Claiming high ecological validity just because it happened outdoors. The task itself has to be something people would realistically do.
Up next: Observations: Watching Behaviour Directly — what happens to your evidence when you stop changing anything and just watch.
Want this explained one-to-one?
Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.