IB Psychology HL Topic 5 — Data Analysis Paper 3 & IA HL only ~11 min read

Averages, Spread and Outliers

Descriptive statistics do two jobs: they tell you where the middle of the data sits, and how far the data spreads out around it. You need both. Two groups can have exactly the same mean and be nothing alike, and one strange value can move the mean somewhere no participant actually was.

📚 What you need to know

The three averages, and what each is for

AverageHow it is foundUse it whenWeakness
MeanAdd all the values, divide by how many there areData are roughly symmetrical with no extreme valuesDragged by outliers; can land on an impossible value
MedianSort the data and take the middle valueThe data are skewed or contain outliersIgnores the actual size of most of the scores
ModeThe most frequently occurring scoreData are categorical, or you need the most common responseMay not exist, or there may be several

The mode is the one students overlook, but it is the only average that works with categorical data. If you have counted how many people chose red, blue or green, a mean colour is meaningless — the mode is all you have. A data set can also have no mode, two modes (bi-modal) or several (multi-modal), and each of those tells you something.

What one outlier does

Here is the damage, made concrete. Take the scores 2, 3, 4, 4, 6, 7 and 9. The mean is 5.0 and the median is 4. Now add one extreme value, 16, and watch what happens.

One value, two very different effects The same seven scores, then the same seven scores plus a 16. WITHOUT THE OUTLIER mean 5.0 median 4 WITH THE OUTLIER 16 outlier mean 6.4 median 5 0 2 4 6 8 10 12 14 16 18 The mean jumps 1.4; the median moves just 1. And the new mean of 6.4 is higher than five of the eight scores in the set.
The mean of 6.4 does not describe anybody. Five of the eight participants scored below it. That is what “the mean is distorted by outliers” actually looks like.
Whenever a question hands you a data set with one obviously odd value, the examiner wants you to notice it, say the median is preferred, and explain that the median is a positional average so extreme values cannot pull it.

Where outliers come from

CauseExampleWhat to do about it
Genuine variabilityTwo people in a sample of 50 have exceptional memoryKeep the data. It is real, and it may be the most interesting part
Novel dataSelf-reported counts, such as how often people check a fitness trackerKeep it, but treat self-report totals cautiously
Collection errorA participant’s score was mistyped or left out of the analysisCorrect it if you can, and say what you did
Never delete an outlier just because it is inconvenient. Removing real data to make a result look tidier is a research integrity issue. Report it, explain it, and use the median instead.

Measuring spread

The range is the simplest measure: highest score minus lowest score. It is quick, but it uses only two values and one outlier ruins it entirely. There is also a convention worth knowing: when data have been rounded, add 1 to the range to allow for the rounding. For 4, 4, 6, 7, 9, 9 that gives 9 − 4 = 5, then 5 + 1 = 6.

Standard deviation is far more informative because it uses every score. It measures how far, on average, scores deviate from the mean. A low standard deviation means scores are tightly clustered, which suggests a consistent, reliable data set. A high one means scores are spread widely, so the mean describes any individual less well.

🧩 Calculating standard deviation, step by step

  1. Calculate the mean of the data set.
  2. Subtract the mean from each score, giving a deviation for every value.
  3. Square each of those deviations, which removes the minus signs.
  4. Add all the squared deviations together.
  5. Divide that total by the number of scores minus 1. This gives the variance.
  6. Take the square root of the variance to get back to the original units. That is the standard deviation.
Why the last step exists Squaring turns the units into units squared.
The square root turns them back into units you can read.

If every participant scored exactly the same — say 15 out of 20 on a memory test — the standard deviation would be zero, because there is no variation at all. The mean, median and mode would all be 15 too. That is a useful edge case to keep in mind, because it shows what standard deviation is really measuring.

Worked examples

WORKED EXAMPLE

Choose the right average and justify it

A researcher records how many words eight participants recalled: 4, 6, 3, 7, 16, 2, 9, 4. Calculate the mean and the median, and state which should be reported.

Step 1: Mean 4 + 6 + 3 + 7 + 16 + 2 + 9 + 4 = 51, and 51 ÷ 8 = 6.375. Step 2: Sort the data for the median 2, 3, 4, 4, 6, 7, 9, 16 Step 3: Median With 8 values, average the 4th and 5th: (4 + 6) ÷ 2 = 5. Step 4: Decide and justify The value 16 is an outlier, well beyond the rest. Report the median of 5, because it is a positional average and is not affected by extreme values. Mean 6.375, median 5; report the median “which should be reported” always wants the justification, not just the number
WORKED EXAMPLE

Interpret two standard deviations

Two classes take the same 20-mark test. Class A has a mean of 14 with a standard deviation of 1.2. Class B has a mean of 14 with a standard deviation of 5.6. Compare the two classes.

Step 1: Compare the central tendency The means are identical, so on average the classes performed the same. Step 2: Compare the dispersion Class A’s standard deviation of 1.2 is much lower, so its scores cluster tightly around 14. Step 3: Describe class B A standard deviation of 5.6 means scores are widely spread, so B contains both much stronger and much weaker performances. Step 4: Draw the conclusion The mean of 14 describes a typical student in A well, and a typical student in B poorly. Reporting the mean alone would have hidden a real difference between the classes. Same mean, very different spread; A is consistent, B is varied this is exactly why every result should report a measure of dispersion

💡 Exam tip

⚠ Common mix-up

Up next: Probability and Inferential Statistics — the point at which you stop describing your data and start asking whether the result was just luck.

Want this explained one-to-one?

Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.

Book a Free Session →