IB Psychology HLTopic 5 — Data AnalysisPaper 3 & IAHL only~11 min read
Averages, Spread and Outliers
Descriptive statistics do two jobs: they tell you where the middle of the data sits, and how far the data spreads out around it. You need both. Two groups can have exactly the same mean and be nothing alike, and one strange value can move the mean somewhere no participant actually was.
📚 What you need to know
Measures of central tendency summarise the typical value: mean, median and mode.
Measures of dispersion summarise the spread: range and standard deviation.
The mean uses every score, which makes it powerful and also fragile.
The median is a positional average and is not affected by extreme values.
A low standard deviation means scores cluster tightly; a high one means they are spread out.
An outlier is a value far beyond the rest, caused by real variability, novel data or a collection error.
When outliers are present, report the median, and say why.
The three averages, and what each is for
Average
How it is found
Use it when
Weakness
Mean
Add all the values, divide by how many there are
Data are roughly symmetrical with no extreme values
Dragged by outliers; can land on an impossible value
Median
Sort the data and take the middle value
The data are skewed or contain outliers
Ignores the actual size of most of the scores
Mode
The most frequently occurring score
Data are categorical, or you need the most common response
May not exist, or there may be several
The mode is the one students overlook, but it is the only average that works with categorical data. If you have counted how many people chose red, blue or green, a mean colour is meaningless — the mode is all you have. A data set can also have no mode, two modes (bi-modal) or several (multi-modal), and each of those tells you something.
What one outlier does
Here is the damage, made concrete. Take the scores 2, 3, 4, 4, 6, 7 and 9. The mean is 5.0 and the median is 4. Now add one extreme value, 16, and watch what happens.
The mean of 6.4 does not describe anybody. Five of the eight participants scored below it. That is what “the mean is distorted by outliers” actually looks like.
Whenever a question hands you a data set with one obviously odd value, the examiner wants you to notice it, say the median is preferred, and explain that the median is a positional average so extreme values cannot pull it.
Where outliers come from
Cause
Example
What to do about it
Genuine variability
Two people in a sample of 50 have exceptional memory
Keep the data. It is real, and it may be the most interesting part
Novel data
Self-reported counts, such as how often people check a fitness tracker
Keep it, but treat self-report totals cautiously
Collection error
A participant’s score was mistyped or left out of the analysis
Correct it if you can, and say what you did
Never delete an outlier just because it is inconvenient. Removing real data to make a result look tidier is a research integrity issue. Report it, explain it, and use the median instead.
Measuring spread
The range is the simplest measure: highest score minus lowest score. It is quick, but it uses only two values and one outlier ruins it entirely. There is also a convention worth knowing: when data have been rounded, add 1 to the range to allow for the rounding. For 4, 4, 6, 7, 9, 9 that gives 9 − 4 = 5, then 5 + 1 = 6.
Standard deviation is far more informative because it uses every score. It measures how far, on average, scores deviate from the mean. A low standard deviation means scores are tightly clustered, which suggests a consistent, reliable data set. A high one means scores are spread widely, so the mean describes any individual less well.
🧩 Calculating standard deviation, step by step
Calculate the mean of the data set.
Subtract the mean from each score, giving a deviation for every value.
Square each of those deviations, which removes the minus signs.
Add all the squared deviations together.
Divide that total by the number of scores minus 1. This gives the variance.
Take the square root of the variance to get back to the original units. That is the standard deviation.
Why the last step exists
Squaring turns the units into units squared.
The square root turns them back into units you can read.
If every participant scored exactly the same — say 15 out of 20 on a memory test — the standard deviation would be zero, because there is no variation at all. The mean, median and mode would all be 15 too. That is a useful edge case to keep in mind, because it shows what standard deviation is really measuring.
Worked examples
WORKED EXAMPLE
Choose the right average and justify it
A researcher records how many words eight participants recalled: 4, 6, 3, 7, 16, 2, 9, 4. Calculate the mean and the median, and state which should be reported.
Step 1: Mean4 + 6 + 3 + 7 + 16 + 2 + 9 + 4 = 51, and 51 ÷ 8 = 6.375.
Step 2: Sort the data for the median2, 3, 4, 4, 6, 7, 9, 16Step 3: Median
With 8 values, average the 4th and 5th: (4 + 6) ÷ 2 = 5.
Step 4: Decide and justify
The value 16 is an outlier, well beyond the rest. Report the median of 5, because it is a positional average and is not affected by extreme values.
Mean 6.375, median 5; report the median“which should be reported” always wants the justification, not just the number
WORKED EXAMPLE
Interpret two standard deviations
Two classes take the same 20-mark test. Class A has a mean of 14 with a standard deviation of 1.2. Class B has a mean of 14 with a standard deviation of 5.6. Compare the two classes.
Step 1: Compare the central tendency
The means are identical, so on average the classes performed the same.
Step 2: Compare the dispersion
Class A’s standard deviation of 1.2 is much lower, so its scores cluster tightly around 14.
Step 3: Describe class B
A standard deviation of 5.6 means scores are widely spread, so B contains both much stronger and much weaker performances.
Step 4: Draw the conclusion
The mean of 14 describes a typical student in A well, and a typical student in B poorly. Reporting the mean alone would have hidden a real difference between the classes.
Same mean, very different spread; A is consistent, B is variedthis is exactly why every result should report a measure of dispersion
💡 Exam tip
Report a measure of central tendency and a measure of dispersion. A mean on its own is an incomplete answer.
Justify your choice of average by referring to skew or outliers, not by preference.
Say the median is a positional average. That word is the reason it resists outliers.
Know the six standard deviation steps in order. The “divide by n minus 1” step is the one people forget.
Link a low standard deviation to consistency and, in a measure, to reliability.
Add 1 to the range when the data have been rounded, and say why.
⚠ Common mix-up
Thinking the mean is always best. It is best for symmetrical data with no outliers, and misleading otherwise.
Believing the median ignores outliers completely. It shifts slightly as the data set grows, it just is not dragged.
Confusing variance and standard deviation. Standard deviation is the square root of the variance.
Reporting a mean with no spread. Two very different data sets can share a mean.
Deleting outliers. Report them and explain them instead.
Confusing dispersion with sample size. A large sample can have tiny spread and vice versa.
Up next: Probability and Inferential Statistics — the point at which you stop describing your data and start asking whether the result was just luck.
Want this explained one-to-one?
Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.