IB Psychology HLTopic 5 — Data AnalysisPaper 3 & IAHL only~11 min read
Choosing the Right Statistical Test
This looks like the hardest thing in the HL course and is actually the most mechanical. Three questions lead you to exactly one test, every single time. Learn the questions in order, learn the grid they produce, and this becomes free marks in both the exam and the IA.
📚 What you need to know
Three questions decide the test: difference or correlation, related or unrelated design, and level of data.
Unrelated means independent measures. Related means repeated measures or matched pairs.
Levels of data: nominal (categories), ordinal (ranked), interval (equal units).
Parametric tests need a normal distribution, interval or ratio data, and homogeneity of variance.
Non-parametric tests make no assumption of normality and work with nominal or ordinal data.
Parametric tests are more powerful, so they are more likely to detect a real effect.
Chi-squared tests both difference and association; Spearman’s rho and Pearson’s r test correlation.
The three questions
Design only matters for tests of difference. A correlation has no conditions to compare, so it skips that branch entirely.
Question 1: difference or correlation?
Look at the hypothesis. If it predicts a difference between conditions, you need a test of difference. If it predicts a relationship between two co-variables, you need a test of correlation. The wording of the hypothesis usually gives it away, which is one more reason to write hypotheses carefully.
Question 2: related or unrelated?
An unrelated design uses independent measures — different people in each condition. A related design uses repeated measures or matched pairs, so the two sets of scores are paired up. This question only applies to tests of difference.
Question 3: what level of data?
Level
What it means
Example
Nominal
Named categories with no order at all
How many people chose yes or no; which of three conditions someone was in
Ordinal
Ranked or ordered, but the gaps are not equal
Rating anxiety 1 to 10; finishing position in a task
Interval
Equal, measurable units between values
Reaction time in milliseconds; number of words recalled
The grid
Level of data
Difference, unrelated design
Difference, related design
Association or correlation
Nominal
Chi-squared
Sign test
Chi-squared
Ordinal
Mann-Whitney U
Wilcoxon T
Spearman’s rho
Interval (parametric)
Unrelated t-test
Related t-test
Pearson’s r
Two details are worth pinning down. Chi-squared appears twice because it tests both difference and association. And Spearman’s rho and Pearson’s r are both tests of correlation — the choice between them is simply whether your data are ordinal or interval.
Say the grid out loud as three columns rather than nine cells. Nominal gives you chi-squared, sign test, chi-squared. Ordinal gives you Mann-Whitney, Wilcoxon, Spearman. Interval gives you the two t-tests and Pearson. It sticks far faster that way.
Parametric and non-parametric
The bottom row of the grid is the parametric row, and parametric tests come with conditions attached. All three must be met.
🧩 The three parametric assumptions
A normal distribution. The data are symmetrical around the mean, with most scores clustered near it, producing the familiar bell curve.
Interval or ratio data. The most sensitive and precise level of measurement, with equal units.
Homogeneity of variance. The conditions have similar dispersion, which you can check by comparing their standard deviations.
Non-parametric tests do not follow these criteria. There is no assumption of a normal distribution, which makes them useful when data are skewed or not continuous — scores on a memory test, for example. They work with nominal or ordinal data and do not depend on homogeneity of variance.
The trade-off is power. Parametric tests are more powerful and precise, meaning they are more likely to detect a significant difference or correlation when one truly exists. So use a parametric test when your data allow it, and a non-parametric test when they do not.
A useful check for homogeneity of variance: if both conditions have similar standard deviations, the data are equally spread and clustered around the mean in each group. Very different standard deviations mean this assumption has failed.
Worked examples
WORKED EXAMPLE
Choose the test and justify it
Twenty participants each complete a memory task twice, once in silence and once with music. The DV is the number of words recalled out of 20. The data are normally distributed and the two conditions have similar standard deviations. Which test should be used?
Step 1: Difference or correlation?
Two conditions are being compared, so this is a test of difference.
Step 2: Related or unrelated?
The same 20 people did both conditions, so this is repeated measures, which is a related design.
Step 3: Level of data?
Number of words recalled has equal units, so the data are interval.
Step 4: Check the parametric assumptions
Normal distribution, interval data and similar standard deviations — all three met, so a parametric test is permitted.
Related t-teststate all three assumptions explicitly when you justify a parametric test
WORKED EXAMPLE
A trickier one: ordinal and unrelated
Two separate groups of participants rate their anxiety on a scale from 1 to 10, one group before a mock exam and one group before a real exam. The distribution is clearly skewed. Which test should be used, and why not a t-test?
Step 1: Difference or correlation?
Two groups are being compared, so a test of difference.
Step 2: Related or unrelated?Two separate groups, so independent measures, which is an unrelated design.
Step 3: Level of data?
A 1 to 10 self-rating is ordinal — the gap between 3 and 4 is not necessarily the same as between 8 and 9.
Step 4: Rule out the parametric option
Two assumptions fail: the data are not interval, and the distribution is skewed rather than normal.
Mann-Whitney Urating scales are ordinal, not interval — this is the classic trap
💡 Exam tip
Answer the three questions in order and write each answer down. That structure alone earns method marks.
Rating scales are ordinal. This single fact decides a large number of exam questions.
List all three parametric assumptions when justifying a t-test or Pearson’s r.
Remember chi-squared appears in two columns, because it tests both difference and association.
Say parametric tests are more powerful and explain what that means: more likely to detect a real effect.
For the IA, name the test and justify it with the three questions. Do not just state it.
⚠ Common mix-up
Treating rating scales as interval. Numbers on a scale do not guarantee equal gaps between them.
Mixing up related and unrelated. Related means the same people or matched pairs; unrelated means separate groups.
Using a t-test on skewed data. Normality is a requirement, not a suggestion.
Confusing Spearman’s rho and Pearson’s r. Spearman for ordinal, Pearson for interval.
Forgetting the sign test. It is the related, nominal option and it is easy to overlook.
Thinking non-parametric means worse. It means fewer assumptions and less power, not lower quality.
Up next: What the Correlation Coefficient Tells You — going deeper into the number itself, and what it can and cannot support.
Want this explained one-to-one?
Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.