IB Psychology HL Topic 5 — Data Analysis Paper 3 & IA HL only ~11 min read

Distributions and the Shape of Data

A distribution is just the shape your data makes when you plot it. That shape decides which average you should trust, whether you are allowed to use a parametric test, and whether an unusual score counts as remarkable. Learn to recognise three shapes and a surprising amount of statistics becomes obvious.

📚 What you need to know

The normal distribution

Height, weight and shoe size all fall into this shape, and so do a lot of psychological measures. Most people sit near the middle, and numbers thin out steadily as you move away in either direction. Because the curve is symmetrical, the three averages land in the same place, which is why the normal distribution is so convenient to work with.

The normal distribution Symmetrical, single peak, thinning tails on both sides. mean = median = mode −2 SD +2 SD the tails never quite touch the axis Beyond two standard deviations counts as extreme. That is how tests for IQ and other traits decide what is unusual.
The tails matter. Because they never touch the axis, the curve refuses to name a final highest or lowest possible score — there is always room for someone more extreme.

Scoring beyond two standard deviations from the mean is what makes a result stand out. That is how an IQ score is judged unusually high or low, how a postpartum depression scale flags a concerning score, and how a low score on an empathy scale becomes clinically interesting. In every case the raw number means nothing until it is placed on the curve.

Normal distributions have a low standard deviation relative to their range, because scores cluster tightly around the mean. That clustering is exactly what the bell shape is showing you.

When the shape goes lopsided

Plenty of real measures are not symmetrical. A very hard exam pushes most scores to the bottom with a few high performers trailing out to the right. A very easy exam does the reverse. Those are skewed distributions, and the direction is named after where the long tail points, not where the bulk of the scores sit.

Skew is named after the tail, not the hump Follow the long tail. That is the direction of the skew. POSITIVE SKEW NEGATIVE SKEW long tail on the right long tail on the left mode median mean The mean always sits furthest into the tail. It is the only average that uses every score, so the extreme values drag it along.
Read the order of the dashed lines from the peak outwards: mode, then median, then mean. That order holds for both skews, just mirrored.
Where the averages land Positive skew: mode < median < mean
Negative skew: mean < median < mode
The naming trips everyone. “Positive skew” sounds like it should mean high scores. It does not — it means most scores are low and the long tail stretches out to the positive end. Look at the tail, never the hump.

What the shape means for your analysis

DistributionReal exampleBest average to reportTest type allowed
NormalHeight, weight, shoe size, IQ in a large sampleMean, because it uses all the data and is not distortedParametric tests are permitted, if the other criteria are met
Positively skewedAge of first job in a population aged 16 to 80; scores on a very hard testMedian, because the long right tail drags the mean upwardsNon-parametric, since the normality assumption fails
Negatively skewedAge of retirement in the same population; scores on a very easy testMedian, because the long left tail drags the mean downwardsNon-parametric, for the same reason

🧩 Identifying the shape from a graph

  1. Find the peak. One peak or two? Two peaks means the data may be two groups mixed together.
  2. Check for symmetry. Fold the curve mentally down the middle. Do the halves match?
  3. Follow the longer tail. Right means positive skew, left means negative skew.
  4. Locate the averages. The mode is at the peak; the mean is furthest into the tail.
  5. Decide what to report. Symmetrical means the mean is safe. Skewed means the median is honest.

Worked examples

WORKED EXAMPLE

Identify the skew and pick the average

A class sits a very easy vocabulary test. Most students score between 18 and 20 out of 20, but three students who were absent for the topic score 4, 5 and 7. Describe the distribution and state which average should be reported.

Step 1: Locate the bulk of the scores Most scores are at the high end, near 18 to 20. Step 2: Locate the long tail The three low scores stretch out to the left, so this is a negative skew. Step 3: Explain what happens to the mean The mean uses every score, so those three low values drag it downwards and it will sit below what a typical student scored. Step 4: Choose the average Report the median, because it is a positional average and is not affected by extreme values. Negatively skewed; report the median “very easy test” is a standard exam signal for negative skew
WORKED EXAMPLE

Work backwards from the averages

For a set of reaction times, the mean is 480 ms, the median is 440 ms and the mode is 420 ms. Describe the shape of the distribution and explain your reasoning.

Step 1: Put the three averages in order mode 420 < median 440 < mean 480 Step 2: Match the order to a shape Mode below median below mean is the signature of a positive skew. Step 3: Describe it in words Most participants reacted quickly, and a small number of very slow reaction times form a long tail to the right. Step 4: Add the implication The mean of 480 ms overstates the typical reaction time, and a non-parametric test would be more appropriate than a parametric one. Positive skew, with a long right tail of slow responses the three averages alone are enough to identify the shape — no graph needed

💡 Exam tip

⚠ Common mix-up

Up next: Averages, Spread and Outliers — the three averages side by side, and exactly how much damage one strange value can do.

Want this explained one-to-one?

Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.

Book a Free Session →