IB Psychology HLTopic 5 — Data AnalysisPaper 3 & IAHL only~10 min read
Frequency Tables and Grouped Data
A frequency table turns a messy pile of scores into something you can actually read. It also contains a trap that catches a very large number of students: when you calculate the mean from a frequency table, you divide by the total frequency, not by the number of rows. Get that wrong and every figure after it is wrong too.
📚 What you need to know
A frequency table records how many times each score, behaviour or category occurred.
A tally is the recording tool; the frequency column is the count written as a number.
The mode is the score with the highest frequency, not the highest score.
The mean is the sum of (score × frequency) divided by the total frequency.
The median is found using a cumulative frequency column, not by eye.
The range is the highest score minus the lowest score in the score column.
Grouped tables use intervals when there are too many distinct values to list.
From raw scores to a summary
Tallying in fives is not decoration — grouping marks into blocks of five is what makes a long count checkable without recounting.
The worked table
Here is a frequency table for the number of goals scored in each match by a school team over one season.
Goals scored
Frequency (number of matches)
Score × frequency
Cumulative frequency
1
1
1
1
2
3
6
4
3
11
33
15
4
20
80
35
5
8
40
43
Total
43
160
43
Now read the four statistics straight off it.
Mode = 4. Four goals happened in 20 matches, more than any other score. Notice the mode is the score in column one, not the frequency of 20.
Mean = 160 ÷ 43 = 3.72. This is the line to be careful with. You add up the score × frequency column to get 160, then divide by the total frequency — the 43 matches. Dividing by 5 because there are five rows would give 32, which is impossible when the highest score is 5. Dividing by 10 would give 16, which is equally impossible. If your mean falls outside the range of your scores, you have divided by the wrong thing.
Median = 4. There are 43 matches, so the middle one is the 22nd. Run down the cumulative frequency column: after score 3 you have reached 15 matches; after score 4 you have reached 35. The 22nd match therefore falls inside the score of 4.
Range = 5 − 1 = 4. Highest score minus lowest score, using the score column.
Mean from a frequency table
mean = sum of (score × frequency) ÷ total frequency
Sanity-check every mean you calculate. It must land somewhere between your lowest and highest score. That one habit catches almost every arithmetic slip on this topic before it costs you marks.
Grouped frequency tables
When there are too many distinct values to list — reaction times in milliseconds, say, or ages across a whole school — the scores are grouped into intervals instead. The trade is straightforward: the table becomes readable, but you lose the individual values, so you can only estimate the mean using the midpoint of each interval.
Interval
Midpoint
Frequency
Midpoint × frequency
0 to 9
4.5
2
9
10 to 19
14.5
6
87
20 to 29
24.5
9
220.5
30 to 39
34.5
3
103.5
Total
—
20
420
The estimated mean is 420 ÷ 20 = 21. It is an estimate because you have assumed every score inside an interval sits exactly at the midpoint, which it will not. The modal class here is 20 to 29 — with grouped data you can name the interval containing the mode, but not the mode itself.
Intervals must not overlap. “0 to 10” followed by “10 to 20” is a genuine error, because a score of 10 could go in either. Use 0 to 9 and 10 to 19, or state the boundary rule explicitly.
🧩 Getting every statistic out of a frequency table
Add a score × frequency column and total it. That total is the sum of all the raw scores.
Total the frequency column. This is your n, and it is what you divide by.
Mean = the first total divided by the second.
Add a cumulative frequency column by running totals down the table.
Median = find position (n + 1) ÷ 2, then read across to the score where the cumulative frequency first passes it.
Mode = the score in the row with the biggest frequency. Range = highest score minus lowest score.
Worked examples
WORKED EXAMPLE
Calculate the mean from a frequency table
A researcher records how many times each of 25 participants checks their phone during a one-hour lesson. Scores of 0, 1, 2, 3 and 4 occur with frequencies of 2, 5, 9, 6 and 3. Calculate the mean, and check it is sensible.
Step 1: Multiply each score by its frequency0×2 = 0, 1×5 = 5, 2×9 = 18, 3×6 = 18, 4×3 = 12Step 2: Total those products0 + 5 + 18 + 18 + 12 = 53Step 3: Total the frequencies2 + 5 + 9 + 6 + 3 = 25, which matches the 25 participants.
Step 4: Divide, then sanity-check53 ÷ 25 = 2.12, which sits between 0 and 4, so it is plausible.
Mean = 2.12 checks per lessonthe frequency total should equal your sample size — a free error check
WORKED EXAMPLE
Find the median using cumulative frequency
Using the same data (scores 0 to 4 with frequencies 2, 5, 9, 6, 3), find the median and the mode.
Step 1: Build the cumulative frequency2, 7, 16, 22, 25Step 2: Find the median position
n = 25, so position = (25 + 1) ÷ 2 = 13th value.
Step 3: Read down the cumulative column
After score 1 you have 7 values; after score 2 you have 16. The 13th value therefore has a score of 2.
Step 4: Mode
The largest frequency is 9, which belongs to a score of 2.
Median = 2, mode = 2the mode is the score, never the frequency itself
💡 Exam tip
Always show the score × frequency column. Method marks live there even if your arithmetic slips.
Divide by the total frequency. Say it out loud as you do it — this is the single most common error on this topic.
Use cumulative frequency for the median rather than trying to count values in your head.
State the modal class for grouped data, and say why you cannot give an exact mode.
Call an estimated mean an estimate, and explain that the midpoint assumption is why.
Check the frequency total equals your stated sample size before you go any further.
⚠ Common mix-up
Dividing by the number of rows. Five rows does not mean n = 5. Divide by the total frequency.
Giving the frequency as the mode. The mode is the score that occurred most often, not how often it occurred.
Reading the median off the middle row. The middle row is not the middle value unless the frequencies happen to be symmetrical.
Calculating the range from the frequency column. The range uses the scores.
Treating a grouped mean as exact. Midpoints are assumptions, so the result is an estimate.
Writing overlapping intervals. Every value must belong to exactly one class.
Up next: Distributions and the Shape of Data — what those frequencies look like when you plot them, and why the shape changes which average you should trust.
Want this explained one-to-one?
Book a free session with an experienced IB Psychology tutor and get your trickiest topics made simple.