IB Psychology HLTopic 5 — Data AnalysisPaper 3 & IAHL only~11 min read
Distributions and the Shape of Data
A distribution is just the shape your data makes when you plot it. That shape decides which average you should trust, whether you are allowed to use a parametric test, and whether an unusual score counts as remarkable. Learn to recognise three shapes and a surprising amount of statistics becomes obvious.
📚 What you need to know
A distribution describes the spread of data around the mean for a sample or population.
A normal distribution is symmetrical, with most scores clustered near the mean, giving the bell curve.
In a perfect normal distribution the mean, median and mode are all at the peak.
The tails never touch the x-axis, because no score is assumed to be the final possible extreme.
A skewed distribution is asymmetrical: one tail is longer than the other.
Positive skew: most scores are low, with a long tail to the right. Negative skew: most scores are high, with a long tail to the left.
The mean is the average most affected by skew, because it uses every score.
The normal distribution
Height, weight and shoe size all fall into this shape, and so do a lot of psychological measures. Most people sit near the middle, and numbers thin out steadily as you move away in either direction. Because the curve is symmetrical, the three averages land in the same place, which is why the normal distribution is so convenient to work with.
The tails matter. Because they never touch the axis, the curve refuses to name a final highest or lowest possible score — there is always room for someone more extreme.
Scoring beyond two standard deviations from the mean is what makes a result stand out. That is how an IQ score is judged unusually high or low, how a postpartum depression scale flags a concerning score, and how a low score on an empathy scale becomes clinically interesting. In every case the raw number means nothing until it is placed on the curve.
Normal distributions have a low standard deviation relative to their range, because scores cluster tightly around the mean. That clustering is exactly what the bell shape is showing you.
When the shape goes lopsided
Plenty of real measures are not symmetrical. A very hard exam pushes most scores to the bottom with a few high performers trailing out to the right. A very easy exam does the reverse. Those are skewed distributions, and the direction is named after where the long tail points, not where the bulk of the scores sit.
Read the order of the dashed lines from the peak outwards: mode, then median, then mean. That order holds for both skews, just mirrored.
Where the averages land
Positive skew: mode < median < mean
Negative skew: mean < median < mode
The naming trips everyone. “Positive skew” sounds like it should mean high scores. It does not — it means most scores are low and the long tail stretches out to the positive end. Look at the tail, never the hump.
What the shape means for your analysis
Distribution
Real example
Best average to report
Test type allowed
Normal
Height, weight, shoe size, IQ in a large sample
Mean, because it uses all the data and is not distorted
Parametric tests are permitted, if the other criteria are met
Positively skewed
Age of first job in a population aged 16 to 80; scores on a very hard test
Median, because the long right tail drags the mean upwards
Non-parametric, since the normality assumption fails
Negatively skewed
Age of retirement in the same population; scores on a very easy test
Median, because the long left tail drags the mean downwards
Non-parametric, for the same reason
🧩 Identifying the shape from a graph
Find the peak. One peak or two? Two peaks means the data may be two groups mixed together.
Check for symmetry. Fold the curve mentally down the middle. Do the halves match?
Follow the longer tail. Right means positive skew, left means negative skew.
Locate the averages. The mode is at the peak; the mean is furthest into the tail.
Decide what to report. Symmetrical means the mean is safe. Skewed means the median is honest.
Worked examples
WORKED EXAMPLE
Identify the skew and pick the average
A class sits a very easy vocabulary test. Most students score between 18 and 20 out of 20, but three students who were absent for the topic score 4, 5 and 7. Describe the distribution and state which average should be reported.
Step 1: Locate the bulk of the scores
Most scores are at the high end, near 18 to 20.
Step 2: Locate the long tail
The three low scores stretch out to the left, so this is a negative skew.
Step 3: Explain what happens to the mean
The mean uses every score, so those three low values drag it downwards and it will sit below what a typical student scored.
Step 4: Choose the average
Report the median, because it is a positional average and is not affected by extreme values.
Negatively skewed; report the median“very easy test” is a standard exam signal for negative skew
WORKED EXAMPLE
Work backwards from the averages
For a set of reaction times, the mean is 480 ms, the median is 440 ms and the mode is 420 ms. Describe the shape of the distribution and explain your reasoning.
Step 1: Put the three averages in ordermode 420 < median 440 < mean 480Step 2: Match the order to a shape
Mode below median below mean is the signature of a positive skew.
Step 3: Describe it in words
Most participants reacted quickly, and a small number of very slow reaction times form a long tail to the right.
Step 4: Add the implication
The mean of 480 ms overstates the typical reaction time, and a non-parametric test would be more appropriate than a parametric one.
Positive skew, with a long right tail of slow responsesthe three averages alone are enough to identify the shape — no graph needed
💡 Exam tip
Say the skew is named after the tail. It is the single sentence that prevents the most common error here.
Learn the average order for each skew. Questions often give you the three averages and nothing else.
Link skew to test choice. Skewed data pushes you towards non-parametric tests, and saying so shows the whole picture.
Mention that the tails never touch the axis if asked to describe a normal distribution fully.
Use two standard deviations as your definition of an extreme score.
Name a real example. Very hard test for positive skew, very easy test for negative, works every time.
⚠ Common mix-up
Thinking positive skew means high scores. It means most scores are low with a long tail to the right.
Believing the median is affected by skew. It shifts a little, but only the mean is genuinely dragged.
Drawing tails that touch the axis. A normal curve approaches the axis without ever reaching it.
Assuming all psychological data is normal. Reaction times, incomes and symptom counts are usually skewed.
Treating skew as an error. It is a genuine property of the data and often the most interesting finding.
Reporting the mean for skewed data without comment. If you report it, say why it may mislead.
Up next: Averages, Spread and Outliers — the three averages side by side, and exactly how much damage one strange value can do.
Want this explained one-to-one?
Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.