IB Psychology SL Topic 5 — Research Design Paper 1 & 2 Core idea ~11 min read

Validity: Measuring What You Meant To

Validity asks a harder question than reliability. Not “would I get this again?” but “is this true, and true of anyone other than the people in the room?” Almost every evaluation point you will ever write in psychology is a validity point wearing a different hat.

📘 What you need to know

The four questions validity asks

When a study reports a result, there are four separate ways it can be wrong, and psychology has a name for each. Learning the names is most of the work.

Four ways a finding can fail to be true VALIDITY INTERNAL EXTERNAL did the IV cause the DV, or did something else? protected by control do the results apply outside this room, sample and year? ecological and temporal CONSTRUCT does it measure the idea it claims? PREDICTIVE does it forecast later behaviour? name the type, then explain it in the words of the study writing only that a study lacks validity is worth almost nothing
Construct and predictive validity sit under the same roof but do not belong to either branch neatly. Treat them as two more questions to ask of any measure.

Internal validity

Internal validity is about the inside of the study. If the noisy group remembered fewer words, was it the noise? Or was it that they were tested at 4pm while the quiet group was tested at 9am? Anything that changes alongside the IV and could explain the result is a confounding variable, and it destroys internal validity.

The defences are the ones you already know: standardised procedures, controlled conditions, random allocation, and using the same materials throughout. When a study is well controlled, the conclusions drawn from it can be trusted, because there is nothing else left to blame.

The strongest phrasing for an exam: “because the two groups differed only in the IV, the difference in the DV can reasonably be attributed to it. This gives the study high internal validity.”

External validity

External validity asks whether the finding travels. Three things can stop it.

Two questions, asked in two different places a study can pass one of these and fail the other badly INSIDE THE STUDY was the effect really caused by the IV? internal validity travels? OUT IN THE WORLD do the findings hold for other people, places and times? external validity tightening control raises one and usually lowers the other good evaluation says which one the study chose to protect
This is the same trade-off you met with lab and field experiments, now with its proper names attached.
A useful sentence to keep: high control buys internal validity and spends external validity. If you can say which one a study bought and which one it sold, you are evaluating properly.

Ecological validity, more carefully

Students often assume ecological validity is only about location. It is really about whether the task resembles something people actually do, and whether their engagement with it is real.

A study can take place in a laboratory and still have decent ecological validity, if participants are genuinely engaged in what they are doing. And a study in a real street can have poor ecological validity if the task is bizarre. Location is a clue, not the answer.

Temporal validity

Some findings age badly, because the world they described has changed. Research on family life carried out when almost every household looked the same will not describe a world of single parents, blended families and same-sex parents. Research on obedience or conformity reflects the social climate it was run in.

Use this carefully. “It is old, so it is invalid” is a weak point. “The social conditions it depended on have changed, so we should not assume the same result today” is a strong one. Name the condition that changed.

Construct and predictive validity

Construct validity matters most for abstract things: intelligence, mood, empathy, depression. You cannot see them, so you build a measure and then have to argue that the measure really captures the idea. A questionnaire that mostly measures how tired someone feels is not a valid measure of depression, however consistent it is.

Predictive validity asks a simpler question: does the score forecast what happens later? A measure of infant attachment has predictive validity if it tells you something about relationships years afterwards. This is what makes psychological measures useful in schools, clinics and workplaces.

TypeThe question it asksThreatened byProtected by
InternalDid the IV cause the DV?Confounding and extraneous variablesControl, standardisation, random allocation
EcologicalIs the task and setting realistic?Artificial tasks and sterile settingsReal settings, meaningful tasks
TemporalDoes it still hold today?Social change since the studyReplication in the present day
ConstructDoes it measure the idea itself?Vague or badly defined conceptsClear definitions, unambiguous tasks
PredictiveDoes it forecast later behaviour?Scores that relate to nothingFollow-up studies over time

Worked examples

WORKED EXAMPLE

Identify the threat to validity

In a study of memory, the quiet condition was run on Monday morning and the noisy condition on Friday afternoon. Identify the problem and explain its effect. [3]

Step 1: Find what changed alongside the IV Time of week and time of day changed with the condition. Step 2: Name it properly A confounding variable — tiredness on Friday afternoon could explain lower recall by itself. Step 3: State the consequence The researcher cannot say the noise caused the difference. Internal validity is reduced the fix is simple: run both conditions at the same time of day, or alternate them
WORKED EXAMPLE

Evaluate the validity of a finding

A lab study of helping behaviour uses 48 first-year psychology students who watch a filmed emergency and press a button if they would help. Evaluate the validity of the conclusions. [6]

Internal validity: strong Controlled conditions and a standardised film mean the IV is the only thing differing between conditions. Ecological validity: weak Pressing a button about a film is not helping anyone. The behaviour measured is not the behaviour of interest. Population validity: weak 48 psychology students are not a cross-section of people, and may already know the theory. Construct validity: questionable Is a button press really a measure of helping, or of stated intention? High internal, low external and construct validity structure the answer by type — it makes the marking obvious

💡 Exam tip

⚠️ Common mix-up

Up next: Generalisability: Who Else Do the Results Apply To? — taking external validity apart and asking exactly how far a finding reaches.

Want this explained one-to-one?

Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.

Book a Free Session →