In review — free for everyone. While a book is in review you are one of its reviewers: read it, use it, and tell us what is wrong. When the reports settle, the Class 10 pass is ₹999 for the year and this book’s PDF is ₹299.
Summarising grouped data
A STORY
Four crates, not two hundred potatoes
Meera lifts a potato from the heap and weighs it in her hand. Arjun drops another into a crate.
“We could weigh every potato,” Arjun says, “but there must be two hundred of them.”
“So we sort them instead,” Meera says. “Under $100$ grams in one crate, $100$ to $200$ in the
next, then $200$ to $300$, then $300$ to $400$.”
“Then we only count how many land in each crate,” Arjun says. “Four numbers, not two hundred.”
“But then we lose each potato’s exact weight,” Meera says. “How do we find the average?”
“Treat every potato in a crate as if it weighed the middle of its range,” Arjun says. “$50$ grams,
$150$, $250$, $350$. The answer comes out close.”
“And the fullest crate tells us where the most common size lies,” Meera says.
Arjun looks along the four crates. “So four counts can still give us the average, the middle and
the most common.”
You will learn to find the mean, the median and the mode when the data comes in groups like these.
Data often arrives already grouped into class intervals, not as single values. A wage survey might report how many workers fall in ₹200–250, then ₹250–300, and so on. Weights, ages and distances get grouped the same way.
This chapter finds the mean, the median and the mode of grouped data. The mean has three methods: direct, assumed-mean and step-deviation. We pick the method that suits the numbers in front of us.
All three methods give the same mean on the same frequency table. The median and the mode each use one formula instead, keyed to one specific class.
Before opening any formula, read the modal class or the median class from the frequency column first. That one step decides which class’s numbers go into the formula. Get this step wrong, and every later number in the solution is wrong too.
Four raw values, 12, 15, 19 and 24, sit inside one class interval, and the class mark 17.5 stands in for all of them.
Grouping keeps all eighteen marks in the count but throws away where each one sat, leaving only five counts to work with.
A SCHOLAR INDIA REMEMBERS
Dignaga was born in south India, around the 5th or 6th century CE. He wrote a book called the Pramanasamuccaya. In it he asked what really counts as a way of knowing. He kept only two. One is what you see. The other is what you work out from what you see. A table of grouped data needs the same care. Some numbers you counted. Others, like a class mark, you worked out.
Dignaga · the kingfisher
Check yourself
This chapter’s methods for the mean, median and mode of grouped data all start from
the individual raw values, listed one by one, exactly as in ungrouped data
a frequency table of class intervals, not the individual raw observations
only the single class with the highest frequency
a graph of the data, plotted before any frequency table is made
Check your answer
the individual raw values, listed one by one, exactly as in ungrouped data — Grouped-data formulas never look at individual values — only the class frequencies, through the class mark.
✓ a frequency table of class intervals, not the individual raw observations — (B) Every formula in this chapter starts from a class’s frequency, not the individual values inside that class.
only the single class with the highest frequency — The modal class alone locates the mode; the mean and median use every class’s frequency, not just the largest one.
a graph of the data, plotted before any frequency table is made — None of the three formulas in this chapter start from a graph — all work directly from the frequency table.
Why can the mean, median and mode of grouped data be computed without knowing the individual values inside each class?
because every value inside a class is actually identical to that class’s mark
because the total number of observations, $n$, is never needed once data is grouped
a class’s frequency is treated as concentrated at one point
because grouped data always has classes of equal size
Check your answer
because every value inside a class is actually identical to that class’s mark — The individual values inside a class are almost never all equal to the class mark — the formula only treats them as if they were, as an approximation.
because the total number of observations, $n$, is never needed once data is grouped — The median formula needs $n$ directly, to find $n/2$ and locate the median class — $n$ is not dropped once data is grouped.
✓ a class’s frequency is treated as concentrated at one point — (C) Grouping trades the individual values for one representative point per class — the class mark for the mean, and the class’s position in the running total for the median and mode.
because grouped data always has classes of equal size — Equal class size affects which formula step simplifies, not why individual values are unnecessary — that comes from using a representative point instead.
A cricket scorer has both the ball-by-ball runs for an innings and the same runs regrouped into class intervals of width 10. What is true of the mean computed each way?
the two means must always come out exactly equal, since grouping never changes the total or the count
the grouped mean is always higher than the ungrouped mean
the grouped mean may differ slightly, since grouping uses each class’s mark
the ungrouped mean cannot be computed once the data has already been grouped
Check your answer
the two means must always come out exactly equal, since grouping never changes the total or the count — Grouping preserves the number of balls, but once each value is replaced by a class mark the sum used for the mean can shift, so the two means need not match.
the grouped mean is always higher than the ungrouped mean — The grouped mean can come out higher or lower than the exact mean, depending on how the values fall inside each class — there is no fixed direction.
✓ the grouped mean may differ slightly, since grouping uses each class’s mark — (C) Grouping approximates every value in a class by that class’s mark, so the grouped mean can come out slightly different from the exact mean of the raw runs.
the ungrouped mean cannot be computed once the data has already been grouped — The ungrouped mean is just the mean of the original raw values — grouping the data afterwards does not erase the ability to compute it, if the raw list survives.
Adding 3 to every value shifts the mean from 4 to 7, and multiplying by 5 scales it from 4 to 20.
You already found the mean and median of a plain list. Now do the same job on grouped data. Try each check below.
Mean of a list (Class 8, Ganita Prakash Part 2, page 107). Here: the direct method finds it by using each class’s mark as $x_i$. Check: The mean of $2, 4, 6$ is $(2 + 4 + 6)/3 = 4$.
Median of a list (Class 8, Ganita Prakash Part 2, page 108). Here: the median class holds the middle-ranked value, once grouped. Check: The median of $1, 3, 5, 7, 9$ is $5$.
Average of two numbers (Class 8, Ganita Prakash Part 2, page 109). Here: the class mark averages a class’s two limits, lower and upper. Check: The average of $8$ and $11$ is $(8 + 11)/2 = 9.5$.
Adding a fixed number (Class 8, Ganita Prakash Part 2, page 106). Here: the assumed-mean method shifts every value by a fixed number, $a$. Check: Adding $3$ to each of $2, 4, 6$ shifts the mean from $4$ to $7$.
A weighted average (Class 8, Ganita Prakash Part 2, page 110). Here: the direct method weights each class mark by its own frequency. Check: Values $4, 6$ with frequencies $3, 2$ average to $4.8$.
Scaling every value (Class 8, Ganita Prakash Part 2, page 108). Here: the step-deviation method undoes this scaling, by a class size $h$. Check: Multiplying each of $2, 4, 6$ by $5$ scales the mean to $20$.
If any of these felt new, read the page named before going on.
The class mark of a class interval is the average of its lower and upper limits. For the interval 10–25, lower limit $10$ and upper limit $25$ give class mark $(10 + 25)/2 = 17.5$.
Every grouped-data formula treats a class’s frequency as sitting at this one point. No formula uses a class’s lower and upper limits on their own. Once we have found the class mark, it is the number we track.
Economics reports household income, and Geography reports rainfall or crop yield, in class intervals such as ₹10,000–₹20,000. Both subjects use the class mark when they compute an average.
Your turn: what is the class mark of the interval 60–75? (Answer: $67.5$, since $(60 + 75)/2 = 67.5$.)
Each class mark sits half a class above its own lower limit, so the marks form a second ruler shifted along from the boundaries.
A SCHOLAR INDIA REMEMBERS
What does the class mark stand for? Say it exactly. It is the middle value of a class. For the class $10$ to $20$, it is $(10 + 20)/2 = 15$. We then treat every value in that class as if it were $15$. That is an assumption. Know that you made it.
Dignaga · the kingfisher
Check yourself
The class mark of the interval 24-36 is
$36 - 24 = 12$
$(24 + 36)/2 = 30$
$24$, the lower limit of the class
$36$, the upper limit of the class
Check your answer
$36 - 24 = 12$ — $36-24=12$ is the class WIDTH, not the class mark — the mark is the midpoint, $(24+36)/2=30$.
✓ $(24 + 36)/2 = 30$ — (B) The class mark averages the two limits: $(24+36)/2 = 30$.
$24$, the lower limit of the class — The class mark averages both limits — using only the lower limit, $24$, ignores half the interval.
$36$, the upper limit of the class — The class mark averages both limits — using only the upper limit, $36$, ignores half the interval.
A table lists the class 45-65 with frequency 8. In the direct-method mean formula, which single number represents every one of those 8 observations?
the frequency, $8$
the class width, $20$
the lower limit, $45$
the class mark, $55$
Check your answer
the frequency, $8$ — $8$ is the frequency — how many observations the class holds — not the value that represents them.
the class width, $20$ — $65-45=20$ is the class width, $h$, used in the median and mode formulas — not the representative value for the mean.
the lower limit, $45$ — The representative value is the midpoint of the class, $55$, not its lower limit alone.
✓ the class mark, $55$ — (D) The direct method treats every observation in a class as if it sat exactly at that class’s mark, $55$.
A shopkeeper records daily sales into the class 150-200. When computing the mean by the direct method, a student uses $150$ instead of the class mark $175$ for this class. What effect does this have?
no effect, since $150$ still lies inside the class 150-200
the computed mean comes out higher than the true grouped mean
the frequency of that class is now wrongly counted twice
the computed mean comes out lower than it should
Check your answer
no effect, since $150$ still lies inside the class 150-200 — Only the midpoint, $175$, is the class’s representative value — using $150$, the lower limit, undercounts that class and does change the mean.
the computed mean comes out higher than the true grouped mean — $150$ is smaller than the true class mark $175$, so this slip pulls the mean down, not up.
the frequency of that class is now wrongly counted twice — The frequency is untouched by this slip — only the value used to represent the class changes, from $175$ to $150$.
✓ the computed mean comes out lower than it should — (D) Using $150$ instead of $175$ undercounts that class’s contribution by $25$ for every one of its observations, so the mean comes out lower than it should.
The modal class is read from one column alone, so a class with big values but a small count can never be it.
The class interval with the highest frequency in a grouped table is the modal class. Nothing more decides it: only the frequency column matters. Scan the frequency column: the modal class is the number with the biggest count.
A distribution can have two classes tied for the highest frequency. That tied case is called bimodal, and it does not appear on these pages.
No other column in the table plays any part in identifying the modal class. Once that one number is found, nothing else on the row matters.
Your turn: classes 0–10, 10–20, 20–30 and 30–40 have frequencies $3, 9, 6, 2$. Which is the modal class? (Answer: $10$–$20$, since $9$ is the highest frequency.)
The tallest bar marks the modal class, with its two neighbouring bars shaded to show the frequencies the mode formula needs.
Check yourself
In a grouped frequency table, the modal class is
the class interval listed first in the table
the class interval containing the largest data values
the middle class interval in the table
the class interval with the highest frequency
Check your answer
the class interval listed first in the table — Table order has nothing to do with it — the modal class is picked by frequency alone, wherever it happens to sit.
the class interval containing the largest data values — The modal class is chosen by how many observations fall in it, not by how large those observations are.
the middle class interval in the table — Table position decides nothing here — only the frequency count decides which class is modal.
✓ the class interval with the highest frequency — (D) The modal class is simply whichever class interval carries the highest frequency in the table.
For the ages of 30 people grouped as 5-15, 15-25, 25-35, 35-45, 45-55 with frequencies 3, 5, 9, 7, 6, the modal class is
25-35
45-55
5-15
35-45
Check your answer
✓ 25-35 — (A) 25-35 carries frequency $9$, the highest in the table, so it is the modal class.
45-55 — 45-55 has frequency $6$, well below the highest frequency, $9$, in 25-35.
5-15 — 5-15 has frequency $3$, the lowest in the table, not the highest.
35-45 — 35-45 has frequency $7$ — higher than most, but still below 25-35’s $9$.
A survey of 50 shoppers’ waiting times gives frequencies 8, 14, 14, 9, 5 across five equal classes. What can be said about the modal class here?
the modal class is simply the first of the two tied classes, by table order
the modal class is whichever tied class has the larger class mark
frequencies can never actually tie in real survey data
two classes tie for the highest frequency; this is bimodal
Check your answer
the modal class is simply the first of the two tied classes, by table order — No rule in this chapter breaks a frequency tie by table order — a genuine tie means the single-modal-class case does not apply.
the modal class is whichever tied class has the larger class mark — Class mark plays no role in choosing between tied classes — a tie simply means this chapter’s one-modal-class formula does not apply.
frequencies can never actually tie in real survey data — Frequencies are whole counts of people, and two classes landing on the same count is entirely possible — this chapter simply does not cover that case.
✓ two classes tie for the highest frequency; this is bimodal — (D) The second and third classes both carry frequency $14$, the highest in the table — whenever two classes tie for the single highest frequency, the data is bimodal and outside the case this chapter’s formula covers.
The median class is the class interval containing the middle-ranked observation. To find it, we add frequencies class by class, in the given order, keeping a running total as we go.
*The median class is the first class where that running total reaches or passes half the total observations, $n/2$.* No later class can be the median class, even if its own frequency is bigger. The class holding the running total the moment it first reaches $n/2$ is the median class.
For example, with $n = 60$, the median class is the first class whose running total reaches at least $30$. The same running total method works on any table, regardless of what $n$ turns out to be.
Your turn: with $n = 80$, what running total first makes a class the median class? (Answer: the first class whose running total reaches $40$, since $n/2 = 80/2 = 40$.)
A filled dot marks where a class starts, an open circle where it ends, so the boundary 20 belongs to only one class.
A SCHOLAR INDIA REMEMBERS
Modal class and median class sound alike. Are they the same kind of thing? No. The modal class has the highest frequency. The median class holds the middle value, and we find it from the cumulative frequency. In one table they can be two different classes.
Dignaga · the kingfisher
Check yourself
The median class of grouped data is located by
the first class where the running total reaches $n/2$
finding the class with the highest frequency
finding the middle class by position in the table, regardless of frequency
finding the class whose class mark is closest to the overall mean
Check your answer
✓ the first class where the running total reaches $n/2$ — (A) Add frequencies class by class, in order, and the median class is the first one where that running total reaches or passes $n/2$.
finding the class with the highest frequency — The highest frequency locates the MODAL class — the median class is found from a running total against $n/2$, which can be a different class entirely.
finding the middle class by position in the table, regardless of frequency — Position in the table means nothing on its own — the median class depends on how the frequencies accumulate, not on which row is physically in the middle.
finding the class whose class mark is closest to the overall mean — The mean plays no part in locating the median class — only the running total of frequencies against $n/2$ does.
68 boxes are weighed and grouped 40-45, 45-50, 50-55, 55-60, 60-65 kg with frequencies 8, 14, 20, 15, 11. With $n/2=34$, the median class is
45-50
50-55
55-60
60-65
Check your answer
45-50 — The running total through 45-50 is only $8+14=22$, still short of $34$ — the median class must be the next one.
✓ 50-55 — (B) The running total reaches $8+14+20=42$, past $34$, at 50-55 — the class before it had only reached $22$.
55-60 — By 55-60 the running total is already $57$ — the running total FIRST passed $34$ one class earlier, at 50-55.
60-65 — 60-65 is where the running total finally reaches the full $68$ — far past the $34$ the median class needs.
A table’s total frequency is $n=41$. What value should each running total be compared against to find the median class?
$n = 41$, the full total
$(n+1)/2 = 21$, the ungrouped-median position formula
$n/2 = 20.5$
$20$, rounding $20.5$ down to a whole number
Check your answer
$n = 41$, the full total — Comparing against the full total $n=41$ would only ever locate the LAST class — the median class needs $n/2$.
$(n+1)/2 = 21$, the ungrouped-median position formula — $(n+1)/2$ gives a POSITION in an ordered list for ungrouped data — the grouped median class is found by comparing the running total to $n/2$ instead.
✓ $n/2 = 20.5$ — (C) The median class is found by comparing the running total to $n/2$, whatever that value comes out to be — here, $20.5$.
$20$, rounding $20.5$ down to a whole number — $n/2$ need not be a whole number — the running total is simply compared against $20.5$ as it stands, with no rounding.
The mean of $n$ observations is their total divided by their count. For observations $x_1, x_2, \dots, x_n$, the mean is $\overline{x} = (\sum x_i)/n$: add everything up, then divide by how many numbers there are.
For example, the mean of $4, 6, 8, 10$ is $(4 + 6 + 8 + 10)/4 = 7$.
Sometimes a value repeats. If a value $x_i$ occurs with frequency $f_i$, the same idea still applies: it reads $\overline{x} = (\sum f_i x_i)/(\sum f_i)$, each value counted as many times as it occurs.
Grouped data only changes what counts as $x_i$, never the total-divided-by-count idea behind the mean.
Your turn: what is the mean of $5, 7, 9, 11$? (Answer: $8$, since $(5 + 7 + 9 + 11)/4 = 32/4 = 8$.)
Eight ungrouped values give a mean of 4.75, a median of 4.5 and a mode of 7, each at a different point.
Check yourself
The mean of the 5 numbers 12, 15, 18, 20, 25 is
$90/5 = 18$
$90$, the sum, left undivided by the count
$18.5$, the average of the smallest and largest values only
$5$, the count of numbers, mistaken for the mean
Check your answer
✓ $90/5 = 18$ — (A) The five numbers sum to $90$; dividing by the count, $5$, gives a mean of $18$.
$90$, the sum, left undivided by the count — $90$ is only the sum — the mean divides that sum by the count of values, $5$, giving $18$.
$18.5$, the average of the smallest and largest values only — Averaging only $12$ and $25$ ignores the three middle values — the mean uses every one of the 5 numbers.
$5$, the count of numbers, mistaken for the mean — $5$ is how many numbers there are, not their mean — the mean is the sum divided by that count.
Why does $\overline{x} = (\sum f_i x_i)/(\sum f_i)$ give the same result as adding every value once and dividing by the count, when a value $x_i$ repeats $f_i$ times?
summing $f_i x_i$ adds $x_i$ to itself $f_i$ times, matching the total
because $f_i$ and $x_i$ always turn out equal for repeated values
because dividing by $\sum f_i$ instead of the raw count gives a more accurate answer
because $f_i x_i$ rounds the value $x_i$ to the nearest frequency
Check your answer
✓ summing $f_i x_i$ adds $x_i$ to itself $f_i$ times, matching the total — (A) $f_i x_i$ stands in for adding $x_i$ to the total $f_i$ separate times, so the weighted sum reaches the same total as writing out every repeated value by hand.
because $f_i$ and $x_i$ always turn out equal for repeated values — $f_i$ (how often a value occurs) and $x_i$ (the value itself) are unrelated numbers — the formula works whether or not they happen to match.
because dividing by $\sum f_i$ instead of the raw count gives a more accurate answer — $\sum f_i$ IS the raw count, once every repeat is counted — the formula is an equivalent way of writing the same mean, not a more accurate one.
because $f_i x_i$ rounds the value $x_i$ to the nearest frequency — $f_i x_i$ is a multiplication, not a rounding step — it scales $x_i$ by how often it occurs.
The value 6 occurs with frequency 3 and the value 4 occurs with frequency 2, and no other value appears. A student computes the mean as $(6+4)/2 = 5$. What is the correct mean, and what did this method get wrong?
$5$, since averaging the two distinct values directly is equivalent
$10$, the sum of the two distinct values
$26$, the weighted sum, left undivided by $n$
$5.2$, weighting each value by its own frequency
Check your answer
$5$, since averaging the two distinct values directly is equivalent — Averaging $6$ and $4$ directly treats them as equally frequent — but $6$ occurs 3 times and $4$ only 2 times, so the true mean is $5.2$, not $5$.
$10$, the sum of the two distinct values — $10$ is just $6+4$ — it ignores both the frequencies and the total count of observations.
$26$, the weighted sum, left undivided by $n$ — $26$ is $\sum f_i x_i$ — dividing by the total count, $n=5$, still needs to happen to reach the mean, $5.2$.
✓ $5.2$, weighting each value by its own frequency — (D) The correct mean is $(6 \cdot 3 + 4 \cdot 2)/5 = 26/5 = 5.2$ — each value must be weighted by how often it occurs, not simply averaged as if it occurred once.
To find the median of ungrouped data, we first arrange the observations in ascending order, smallest to largest. The median is the value sitting at the middle position.
If the count $n$ is odd, one value sits exactly in the middle: the median is the value at position $(n + 1)/2$. An odd count always leaves exactly one middle value.
If $n$ is even, two values share the middle. *The median is then the average of the values at positions $n/2$ and $n/2 + 1$.* An even count always averages the two values sitting either side of the middle.
For example, the ordered data $3, 7, 9, 12$ has $n = 4$, even. The median averages positions 2 and 3: $(7 + 9)/2 = 8$.
Your turn: what is the median of $2, 5, 6, 8, 9$? (Answer: $6$, since $n = 5$ is odd and $6$ sits at position $(5 + 1)/2 = 3$.)
The median averages the middle two values, 7 and 9, to 8 for four values, or is the middle value, 6, for five.
Check yourself
Arranged in order, 8, 11, 15, 19, 23 has $n=5$ (odd). The median is
$8$, the smallest value
$23$, the largest value
$(8+23)/2 = 15.5$, the average of the smallest and largest
$15$, the value at position $(5+1)/2 = 3$
Check your answer
$8$, the smallest value — $8$ is the smallest value in the list, not the one at the middle position, $15$.
$23$, the largest value — $23$ is the largest value in the list, not the one at the middle position, $15$.
$(8+23)/2 = 15.5$, the average of the smallest and largest — Averaging the extremes, $8$ and $23$, gives $15.5$ — close to, but not the same as, the true middle value, $15$.
✓ $15$, the value at position $(5+1)/2 = 3$ — (D) With $n=5$ odd, the median sits at position $(5+1)/2=3$, which is $15$.
For 9, 13, 17, 25, 31, 35 (already in order, $n=6$, even), the median is
$17$, the value at position $n/2=3$ alone
$25$, the value at position $n/2+1=4$ alone
$(17+25)/2 = 21$
$(9+35)/2 = 22$, the average of the smallest and largest
Check your answer
$17$, the value at position $n/2=3$ alone — For even $n$, the median averages BOTH middle positions, $3$ and $4$ — using $17$ alone drops the second one.
$25$, the value at position $n/2+1=4$ alone — For even $n$, the median averages BOTH middle positions, $3$ and $4$ — using $25$ alone drops the first one.
✓ $(17+25)/2 = 21$ — (C) With $n=6$ even, the median averages the two middle values, at positions $3$ and $4$: $(17+25)/2=21$.
$(9+35)/2 = 22$, the average of the smallest and largest — Averaging the extremes, $9$ and $35$, gives $22$ — close to, but not the same as, averaging the two true middle values, $21$.
A data set has $n=7$ values sorted in ascending order. Which position holds the median?
position $7$, the last value in the list
position $(7+1)/2=4$, the single middle value
the average of the values at positions 3 and 4, as if $n$ were even
position $7/2 = 3.5$, left as a fractional position
Check your answer
position $7$, the last value in the list — Position $7$ is the last value in the sorted list, not the middle one, position $4$.
✓ position $(7+1)/2=4$, the single middle value — (B) With $n=7$ odd, there is one exact middle position, $(7+1)/2=4$ — no averaging is needed.
the average of the values at positions 3 and 4, as if $n$ were even — Averaging two positions is only for EVEN $n$ — with $n=7$ odd, there is a single exact middle position, $4$, and no averaging is needed.
position $7/2 = 3.5$, left as a fractional position — A position in an ordered list must be a whole number — the correct formula for odd $n$, $(n+1)/2$, always gives one directly.
The mode of ungrouped data is the value that occurs most often: the observation with the highest frequency. We count how many times each value shows up, and the mode is the value with the highest count.
For example, in the data $2, 3, 3, 5, 3, 7$, the value $3$ occurs three times: more than any other value. So the mode is $3$.
Unlike the mean or the median, the mode need not be unique. Data with two equally frequent values is called bimodal. Bimodal data does not appear on these pages, but the name is worth knowing for when two values tie.
Your turn: what is the mode of $4, 6, 6, 9, 6, 2$? (Answer: $6$, since it occurs three times, more than any other value.)
Check yourself
For the values 7, 3, 7, 9, 7, 3, 5, the mode is
$3$, since it is the smallest value that repeats
$9$, the largest value in the list
$7$, the most frequent value
the mean of all the values
Check your answer
$3$, since it is the smallest value that repeats — $3$ occurs only twice — fewer times than $7$, which occurs 3 times and is the true mode.
$9$, the largest value in the list — $9$ occurs only once — the mode is about frequency, not size, and $7$ occurs most often.
✓ $7$, the most frequent value — (C) $7$ occurs 3 times — more often than $3$ (twice) or any other value (once each) — so $7$ is the mode.
the mean of all the values — The mode is the most frequent value; the mean is the sum of all values divided by their count — the two measures answer different questions.
Data can be bimodal, but this chapter’s mode formula does not extend to that case. Why not?
because bimodal data cannot occur in real measurements
because the formula always averages the two most frequent values instead
because bimodal data has no mean or median either
the formula needs one class with the strictly highest frequency
Check your answer
because bimodal data cannot occur in real measurements — Two values tying for the highest frequency is entirely possible in real data — this chapter simply does not extend its formula to that case.
because the formula always averages the two most frequent values instead — The mode formula has no built-in way to average two tied values — it is written for exactly one most-frequent value.
because bimodal data has no mean or median either — A frequency tie only affects the MODE — the mean and median can still be computed from bimodal data without any trouble.
✓ the formula needs one class with the strictly highest frequency — (D) The mode formula picks out the ONE most frequent value — when two values tie for the highest frequency, there is no single value left for it to identify.
A shop’s shoe sales are: size 7 sold 12 times, size 8 sold 12 times, size 9 sold 6 times. A clerk reports the mode as size 8, because it is listed later than size 7 in the table. What is wrong with this reasoning?
sizes 7 and 8 tie for the highest frequency; the data is bimodal
nothing is wrong — later position in a table always breaks a frequency tie
size 9 is actually the mode, since it is the last size listed with any sales
the true mode is the average of sizes 7 and 8, giving $7.5$
Check your answer
✓ sizes 7 and 8 tie for the highest frequency; the data is bimodal — (A) Both size 7 and size 8 sell $12$ times each, tied for the highest — the data is bimodal, and no rule in this chapter breaks such a tie by table order.
nothing is wrong — later position in a table always breaks a frequency tie — No rule breaks a tie by table position — sizes 7 and 8, both at frequency $12$, are genuinely tied, making the data bimodal.
size 9 is actually the mode, since it is the last size listed with any sales — Size 9 sold only $6$ times, the LOWEST count of the three — table position plays no role in identifying the mode.
the true mode is the average of sizes 7 and 8, giving $7.5$ — There is no rule that averages two tied values into a single mode — with sizes 7 and 8 tied at frequency $12$, the data is simply bimodal.
Three classes stand at their class marks, 5, 15 and 25, weighted 3, 4 and 3, giving a mean of 15.
The direct method finds the mean of grouped data straight from the class marks and frequencies. The formula is $\overline{x} = (\sum f_i x_i)/(\sum f_i)$, summed over every class in the table. Every number this formula needs is already sitting in the frequency table.
Each class’s own class mark serves as its $x_i$, paired with that class’s own frequency $f_i$. We build a small table, class mark next to frequency, before multiplying a single pair.
Use the class mark for $x_i$, never a class’s boundary numbers on their own. If we ever reach for the boundary numbers instead, the fix is to stop and go back to the class mark.
Your turn: classes 0–10, 10–20 and 20–30 have frequencies $3, 4, 3$ and class marks $5, 15, 25$. What is the mean by the direct method? (Answer: $15$, since the products add to $150$ and the frequencies add to $10$, so $150/10 = 15$.)
Check yourself
By the direct method, the mean of grouped data is computed as
$\overline{x} = (\sum f_i)/(\sum x_i)$, with frequencies and marks swapped
$\overline{x} = (\sum x_i)/n$, ignoring the frequencies entirely
$\overline{x} = (\sum f_i x_i)/(\sum f_i)$, using each class’s mark $x_i$
$\overline{x} = (\sum f_i l_i)/(\sum f_i)$, using each class’s lower limit instead of its mark
Check your answer
$\overline{x} = (\sum f_i)/(\sum x_i)$, with frequencies and marks swapped — The weighted sum $\sum f_i x_i$ belongs in the numerator, and the total frequency $\sum f_i$ in the denominator — swapping them gives a different, wrong quantity.
$\overline{x} = (\sum x_i)/n$, ignoring the frequencies entirely — Without weighting by $f_i$, every class mark counts once regardless of how many observations actually fall in it — the frequencies cannot be dropped.
✓ $\overline{x} = (\sum f_i x_i)/(\sum f_i)$, using each class’s mark $x_i$ — (C) The direct method weights each class mark $x_i$ by its frequency $f_i$, sums those products, and divides by the total frequency.
$\overline{x} = (\sum f_i l_i)/(\sum f_i)$, using each class’s lower limit instead of its mark — The direct method’s representative value is the class MARK, the midpoint — not the lower limit, which always understates it.
The runs of 25 batsmen are grouped as 0-20, 20-40, 40-60, 60-80, 80-100 with frequencies 3, 7, 9, 4, 2. By the direct method, the mean is
$36$, from using each class’s lower limit instead of its mark
$1150$, the value of $\sum f_i x_i$, left undivided
$50$, the class mark of the middle class alone
$46$
Check your answer
$36$, from using each class’s lower limit instead of its mark — Using the lower limits ($0,20,40,60,80$) instead of the marks gives $900/25=36$ — undercounting every class by half its width.
$1150$, the value of $\sum f_i x_i$, left undivided — $1150$ is only $\sum f_i x_i$ — dividing by $\sum f_i=25$ still needs to happen, giving $46$.
$50$, the class mark of the middle class alone — $50$ is just the mark of the middle class, 40-60 — the mean weights every class by its own frequency, not one class alone.
Why does the direct method use the class mark $x_i$ rather than the lower limit of each class, when computing $\sum f_i x_i$?
because the class mark is always a whole number, unlike the lower limit
the class mark is the midpoint; the lower limit always understates the class
because the lower limit is reserved for the median formula and never appears in a mean calculation
because using the lower limit would make $\sum f_i$ come out wrong
Check your answer
because the class mark is always a whole number, unlike the lower limit — A class mark is a whole number only when the limits happen to add to an even number — it is chosen for being the midpoint, not for being a whole number.
✓ the class mark is the midpoint; the lower limit always understates the class — (B) The mark sits at the midpoint of the class, balancing the values above and below it — the lower limit always sits below every value in the class, biasing the mean downward.
because the lower limit is reserved for the median formula and never appears in a mean calculation — The lower limit does appear in the median and mode formulas, but that is not the reason the direct-method mean avoids it — the mean needs a representative value, and the mark fits that role better.
because using the lower limit would make $\sum f_i$ come out wrong — $\sum f_i$ is just the total frequency — it does not depend on which representative value, mark or limit, is chosen for $x_i$.
50 packets of chips are grouped as 20-30, 30-40, 40-50, 50-60 g with frequencies 10, 15, 20, 5. By the direct method, the mean weight is
$34$ g, from using lower limits instead of class marks
$1950$ g, the value of $\sum f_i x_i$, left undivided
$39$ g
$487.5$ g, from dividing by the number of classes (4) instead of the total frequency (50)
Check your answer
$34$ g, from using lower limits instead of class marks — Using the lower limits ($20,30,40,50$) instead of the marks gives $1700/50=34$ g, undercounting every class.
$1950$ g, the value of $\sum f_i x_i$, left undivided — $1950$ is only $\sum f_i x_i$ — dividing by $\sum f_i=50$ still needs to happen, giving $39$ g.
✓ $39$ g — (C) $\sum f_i x_i = 10(25)+15(35)+20(45)+5(55) = 1950$, and $\sum f_i = 50$, so $\overline{x} = 1950/50 = 39$ g.
$487.5$ g, from dividing by the number of classes (4) instead of the total frequency (50) — The denominator must be $\sum f_i=50$, the total number of packets — dividing by the number of classes, $4$, gives an inflated, meaningless figure.
The same fifty packets are measured on two rulers, one in grams and one in deviations, with the mean at the same spot on both.
When the class marks are large, working with them directly gets clumsy: it means multiplying big numbers for no real gain. The assumed-mean method picks one class mark as an assumed mean, $a$, and works with deviations instead.
Set $d_i = x_i - a$ for every class. Then $\overline{x} = a + (\sum f_i d_i)/(\sum f_i)$: the same mean as the direct method, from smaller numbers. We pick $a$ near the middle of the class marks, and the deviations stay small on both sides.
*Any valid choice of $a$ gives exactly the same mean.* On the tea-packet table, $a = 150$ and $a = 110$ both give $150.8$ grams.
Your turn: classes with marks $10, 20, 30$ and frequencies $2, 5, 3$ take $a = 20$. What is the mean? (Answer: $21$, since the deviations are $-10, 0, 10$, their weighted total is $10$, and $20 + 10/10 = 21$.)
Check yourself
The assumed-mean method rewrites the mean as
$\overline{x} = a - (\sum f_i d_i)/(\sum f_i)$, subtracting the correction instead of adding it
$\overline{x} = a + (\sum d_i)/(\sum f_i)$, leaving the frequency out of the numerator
$\overline{x} = a + (\sum f_i d_i)/(\sum f_i)$, where $d_i = x_i - a$
Check your answer
$\overline{x} = a - (\sum f_i d_i)/(\sum f_i)$, subtracting the correction instead of adding it — The correction term is ADDED back to $a$, not subtracted — subtracting it would push the mean the wrong way.
$\overline{x} = (\sum f_i d_i)/(\sum f_i)$, dropping $a$ entirely — Without adding $a$ back, the formula only gives the correction, not the actual mean — $a$ must be added on at the end.
$\overline{x} = a + (\sum d_i)/(\sum f_i)$, leaving the frequency out of the numerator — The numerator must be the FREQUENCY-weighted sum $\sum f_i d_i$, not the plain sum of deviations $\sum d_i$.
✓ $\overline{x} = a + (\sum f_i d_i)/(\sum f_i)$, where $d_i = x_i - a$ — (D) A chosen assumed mean $a$ is corrected by adding back the weighted average deviation, $(\sum f_i d_i)/(\sum f_i)$.
Daily wages of 40 workers are grouped 100-120, 120-140, 140-160, 160-180, 180-200 with frequencies 5, 9, 15, 7, 4. Taking $a=150$ gives deviations $d_i=-40,-20,0,20,40$ and $\sum f_i d_i=-80$. The mean is
$150$, just the assumed mean itself
$148$
$-2$, the correction term alone, without adding back $a$
$152$, from adding the correction with the wrong sign
Check your answer
$150$, just the assumed mean itself — $150$ is only the assumed mean, $a$ — the correction, $-2$, still needs to be added to reach $148$.
$-2$, the correction term alone, without adding back $a$ — $-2$ is just $(\sum f_i d_i)/(\sum f_i)$ — adding it back to $a=150$ gives the actual mean, $148$.
$152$, from adding the correction with the wrong sign — The correction here is $-2$, not $+2$ — adding $150+(-2)$ gives $148$, not $150+2=152$.
Why is $d_i = x_i - a$ used instead of $x_i$ itself in the assumed-mean method?
because $x_i$ is not permitted to appear anywhere in a mean formula, only deviations are
because using $d_i$ removes the need to know the frequencies at all
because $a$ must always equal the true mean for the method to work
$d_i$ is a smaller number than $x_i$, easier to sum by hand
Check your answer
because $x_i$ is not permitted to appear anywhere in a mean formula, only deviations are — The direct method uses $x_i$ directly, with no such restriction — the assumed-mean method only shrinks the numbers involved, for convenience.
because using $d_i$ removes the need to know the frequencies at all — The frequencies $f_i$ are still needed to weight each $d_i$ in $\sum f_i d_i$ — switching to deviations shrinks the numbers, not the need for frequencies.
because $a$ must always equal the true mean for the method to work — $a$ can be any convenient value, not necessarily the true mean — the method corrects for whatever gap exists between $a$ and the true mean.
✓ $d_i$ is a smaller number than $x_i$, easier to sum by hand — (D) Subtracting a nearby $a$ shrinks every class mark down to a small deviation, making the weighted sum far easier to compute by hand.
For the same wage data, taking $a=130$ instead gives deviations $d_i=-20,0,20,40,60$ with $\sum f_i d_i=720$ over $\sum f_i=40$. What mean does this give?
$130$, the new assumed mean, unadjusted
$18$, the correction term alone, without adding back $a$
$850$, adding $a$ and $\sum f_i d_i$ without ever dividing by the total frequency
$148$, the same mean as with $a=150$
Check your answer
$130$, the new assumed mean, unadjusted — $130$ is only the assumed mean — the correction, $18$, still needs to be added to reach $148$.
$18$, the correction term alone, without adding back $a$ — $18$ is just $720/40$ — adding it back to $a=130$ gives the actual mean, $148$.
$850$, adding $a$ and $\sum f_i d_i$ without ever dividing by the total frequency — $720$ must be divided by $\sum f_i=40$ before adding it to $a$ — adding it raw gives $130+720=850$, far too large.
✓ $148$, the same mean as with $a=150$ — (D) $\overline{x} = 130 + 720/40 = 130+18=148$ — the same mean as before, whichever valid $a$ is chosen.
Sometimes every deviation $d_i$ shares a common factor: the class size, $h$. The step-deviation method divides that factor out before multiplying, so the numbers to multiply are far smaller.
Set $u_i = (x_i - a)/h$ for every class. Then $\overline{x} = a + h \cdot (\sum f_i u_i)/(\sum f_i)$: the smallest numbers of the three methods to multiply by hand. We check the deviations for a shared factor before dividing: it is not always there.
Otherwise, the assumed-mean method is just as easy.
Your turn: classes with marks $20, 40, 60$ and frequencies $3, 7, 2$ take $a = 40$ and $h = 20$. What is the mean? (Answer: $\approx 38.33$, since $u_i = -1, 0, 1$, their weighted total is $-1$, and $40 + 20 \cdot (-1/12) \approx 38.33$.)
Check yourself
The step-deviation method sets $u_i = (x_i-a)/h$ and computes the mean as
$\overline{x} = a + (\sum f_i u_i)/(\sum f_i)$, without multiplying by $h$ at the end
$\overline{x} = a \cdot h + (\sum f_i u_i)/(\sum f_i)$, multiplying $a$ by $h$ instead of the correction term
$\overline{x} = h + a \cdot (\sum f_i u_i)/(\sum f_i)$, with the roles of $a$ and $h$ swapped
$\overline{x} = a + h \cdot (\sum f_i u_i)/(\sum f_i)$
Check your answer
$\overline{x} = a + (\sum f_i u_i)/(\sum f_i)$, without multiplying by $h$ at the end — Dividing by $h$ to form $u_i$ must be undone by multiplying the correction term back by $h$ — skipping that step leaves the mean too close to $a$.
$\overline{x} = a \cdot h + (\sum f_i u_i)/(\sum f_i)$, multiplying $a$ by $h$ instead of the correction term — It is the correction term, $(\sum f_i u_i)/(\sum f_i)$, that gets multiplied by $h$ — not the assumed mean $a$ itself.
$\overline{x} = h + a \cdot (\sum f_i u_i)/(\sum f_i)$, with the roles of $a$ and $h$ swapped — $a$ is the base value added plainly; $h$ is the class size that scales the correction term back up — swapping their roles breaks the formula.
✓ $\overline{x} = a + h \cdot (\sum f_i u_i)/(\sum f_i)$ — (D) After dividing every deviation by $h$, the correction term must be multiplied back by $h$ before adding it to $a$.
Monthly bills (in hundreds of rupees) for 68 households are grouped 0-20, 20-40, 40-60, 60-80, 80-100. Taking $a=50$, $h=20$ gives $u_i=-2,-1,0,1,2$ and $\sum f_i u_i=34$. The mean is
$60$
$50.5$, from leaving out the multiplication by $h$
$50$, the assumed mean before any correction
$730$, adding $a$ to $h$ times the raw total, skipping the division by $\sum f_i$
$50.5$, from leaving out the multiplication by $h$ — $50+34/68=50.5$ leaves out multiplying the correction by $h=20$ — doing so gives the true mean, $60$.
$50$, the assumed mean before any correction — $50$ is only the assumed mean — the correction, $10$, still needs to be added to reach $60$.
$730$, adding $a$ to $h$ times the raw total, skipping the division by $\sum f_i$ — $20 \cdot 34=680$ must be divided by $\sum f_i=68$ before adding it to $a$ — adding it raw gives $50+680=730$, far too large.
Step-deviation divides out $h$ from every deviation. When is this shortcut actually a saving over the assumed-mean method?
always, whatever the class sizes happen to be
only when the assumed mean $a$ equals the true mean
when every class shares the same class size $h$
only when the total frequency $\sum f_i$ is an exact multiple of $h$
Check your answer
always, whatever the class sizes happen to be — If class sizes differ, dividing by a single $h$ does not give clean whole-number $u_i$ values — the saving only holds when every class shares the same size.
only when the assumed mean $a$ equals the true mean — $a$ can be any convenient value, whether or not it equals the true mean — the shortcut depends on equal class sizes, not on how close $a$ is to the mean.
✓ when every class shares the same class size $h$ — (C) Equal class sizes make every deviation a clean multiple of $h$, so dividing by $h$ leaves small whole numbers that are easy to multiply by hand.
only when the total frequency $\sum f_i$ is an exact multiple of $h$ — The shortcut’s saving comes from equal class sizes making $u_i$ small whole numbers — $\sum f_i$ being a multiple of $h$ has nothing to do with it.
A different table has $h=10$, $a=45$, and $\sum f_i u_i=-18$ over $\sum f_i=60$. What mean does this give?
$45$, the bare assumed mean, no correction applied
$44.7$, from leaving out the multiplication by $h$
$42$
$-135$, adding $a$ to $h$ times the raw total, never divided by $\sum f_i$
Check your answer
$45$, the bare assumed mean, no correction applied — $45$ is only the assumed mean — the correction, $-3$, still needs to be added to reach $42$.
$44.7$, from leaving out the multiplication by $h$ — $45+(-18/60)=44.7$ leaves out multiplying the correction by $h=10$ — doing so gives the true mean, $42$.
$-135$, adding $a$ to $h$ times the raw total, never divided by $\sum f_i$ — $10 \cdot (-18)=-180$ must be divided by $\sum f_i=60$ before adding it to $a$ — adding it raw gives $45-180=-135$, wildly off.
$l + ((f_1-f_2)/(2f_1-f_0-f_2)) \cdot h$, with $f_0$ and $f_2$ swapped in the numerator
$((f_1-f_0)/(2f_1-f_0-f_2)) \cdot h$, without adding $l$
$l + ((f_1-f_0)/(f_1-f_0-f_2)) \cdot h$, missing the factor of 2 on $f_1$ in the denominator
$l + ((f_1-f_0)/(2f_1-f_0-f_2)) \cdot h$, for the modal class
Check your answer
$l + ((f_1-f_2)/(2f_1-f_0-f_2)) \cdot h$, with $f_0$ and $f_2$ swapped in the numerator — The numerator must be $f_1-f_0$, using the class BEFORE the modal class — swapping in $f_2$ there gives a different formula.
$((f_1-f_0)/(2f_1-f_0-f_2)) \cdot h$, without adding $l$ — The fraction only gives a correction distance INTO the modal class — $l$, the class’s lower limit, must be added to locate the mode itself.
$l + ((f_1-f_0)/(f_1-f_0-f_2)) \cdot h$, missing the factor of 2 on $f_1$ in the denominator — The denominator is $2f_1-f_0-f_2$, with $f_1$ doubled — dropping that factor of 2 changes the formula’s value.
✓ $l + ((f_1-f_0)/(2f_1-f_0-f_2)) \cdot h$, for the modal class — (D) $l$ and $h$ come from the modal class itself; $f_1$ is its frequency, and $f_0$, $f_2$ its two immediate neighbours.
For ages grouped 5-15, 15-25, 25-35, 35-45, 45-55 with frequencies 3, 5, 9, 7, 6, the modal class is 25-35 with $l=25$, $h=10$, $f_1=9$, $f_0=5$, $f_2=7$. The mode is
$28.33$, from swapping $f_0$ and $f_2$ in the formula
$6.67$, from leaving out $l$
$30$, the class mark of the modal class, mistaken for the mode
$31.67$
Check your answer
$28.33$, from swapping $f_0$ and $f_2$ in the formula — Swapping the neighbours gives $25+((9-7)/(18-7-5)) \cdot 10 = 28.33$ — the correct reading, $f_0=5$ and $f_2=7$, gives $31.67$.
$6.67$, from leaving out $l$ — $(4/6) \cdot 10=6.67$ is only the correction distance — adding $l=25$ gives the actual mode, $31.67$.
$30$, the class mark of the modal class, mistaken for the mode — $30$ is the midpoint of 25-35 — the mode formula shifts away from that midpoint toward whichever neighbour has the closer frequency, landing on $31.67$.
Why does the mode formula subtract both $f_0$ and $f_2$ from $2f_1$ in the denominator, rather than using $f_1$ alone?
the subtraction is only there to keep the denominator from being zero
using $f_1$ alone in the denominator would give exactly the same mode
it compares the modal class’s frequency against both neighbours at once
the denominator’s only job is to convert the class size $h$ into the same units as the frequencies
Check your answer
the subtraction is only there to keep the denominator from being zero — The denominator can still come out small or, rarely, zero — its role is to compare $f_1$ against both neighbours, not to avoid a specific value.
using $f_1$ alone in the denominator would give exactly the same mode — Dropping $f_0$ and $f_2$ from the denominator changes its value and so changes the computed mode — they are not optional terms.
✓ it compares the modal class’s frequency against both neighbours at once — (C) Both neighbouring frequencies enter the denominator so the formula can weigh which side of the modal class the data leans toward.
the denominator’s only job is to convert the class size $h$ into the same units as the frequencies — Frequencies are plain counts with no unit to convert — the denominator’s role is comparing $f_1$ to its neighbours, and $h$ is applied separately, outside the fraction.
Marks are grouped 0-10, 10-20, 20-30, 30-40, 40-50 with frequencies 4, 12, 20, 9, 5. The modal class is 20-30, with $l=20$, $h=10$, $f_1=20$, $f_0=12$, $f_2=9$. The mode is
$25.79$, from swapping $f_0$ and $f_2$
$4.21$, from leaving out $l$
$25$, the class mark of the modal class, mistaken for the mode
$24.21$
Check your answer
$25.79$, from swapping $f_0$ and $f_2$ — Swapping the neighbours gives $20+((20-9)/(40-9-12)) \cdot 10=25.79$ — the correct reading, $f_0=12$ and $f_2=9$, gives $24.21$.
$4.21$, from leaving out $l$ — $(8/19) \cdot 10=4.21$ is only the correction distance — adding $l=20$ gives the actual mode, $24.21$.
$25$, the class mark of the modal class, mistaken for the mode — $25$ is the midpoint of 20-30 — the mode formula shifts away from that midpoint toward the nearer-frequency neighbour, landing on $24.21$.
Nothing in the mode formula is worked out from scratch: all five inputs are read straight off three neighbouring rows of the table.
The mode of grouped data comes from one formula, applied to the modal class. It is $l + ((f_1 - f_0)/(2 f_1 - f_0 - f_2)) \cdot h$. Five numbers go into this one formula, and every one of them comes from the frequency table.
Here $l$ is the modal class’s lower limit and $h$ is the class size. $f_1$ is the modal class’s own frequency. $f_0$ and $f_2$ are its two neighbours’ frequencies: the classes just before and after it. We find these five numbers on the table before substituting a single one into the formula.
Get $f_0$ and $f_2$ right and the final substitution is just arithmetic.
For example, take $f_1 = 8$, $f_0 = 7$ and $f_2 = 2$, with class size $2$ and lower limit $3$. Then mode $= 3 + ((8 - 7)/(16 - 7 - 2)) \cdot 2 \approx 3.286$.
Your turn: a modal class has $f_1 = 15$, $f_0 = 9$, $f_2 = 6$, class size $4$ and lower limit $20$. What is the mode? (Answer: $21.6$, since $20 + ((15 - 9)/(30 - 9 - 6)) \cdot 4 = 20 + 1.6 = 21.6$.)
$l + ((n-C)/f) \cdot h$, using $n$ instead of $n/2$
$l + ((n/2-C)/f) \cdot h$, for the median class
$((n/2-C)/f) \cdot h$, without adding $l$
$l + ((n/2-C)/h) \cdot f$, with $f$ and $h$ swapped
Check your answer
$l + ((n-C)/f) \cdot h$, using $n$ instead of $n/2$ — The formula compares against $n/2$, half the total — using the full $n$ gives a much larger, wrong correction.
✓ $l + ((n/2-C)/f) \cdot h$, for the median class — (B) $l$, $f$ and $h$ come from the median class itself; $C$ is the running total of frequencies BEFORE that class.
$((n/2-C)/f) \cdot h$, without adding $l$ — The fraction only gives a correction distance INTO the median class — $l$, the class’s lower limit, must be added to locate the median itself.
$l + ((n/2-C)/h) \cdot f$, with $f$ and $h$ swapped — $f$, the median class’s own frequency, belongs in the denominator, and $h$, the class size, multiplies the whole fraction — swapping them changes the formula’s value.
68 boxes are grouped 40-45, 45-50, 50-55, 55-60, 60-65 kg with frequencies 8, 14, 20, 15, 11. Here $n/2=34$, the median class is 50-55, with $C=22$ (the total before it) and $f=20$. The median is
$61.5$ kg, from using $n=68$ instead of $n/2=34$
$48$ kg, from using the running total through 50-55 itself, $C=42$, instead of the total before it
$53$ kg
$50$ kg, the lower limit of the median class, mistaken for the median
Check your answer
$61.5$ kg, from using $n=68$ instead of $n/2=34$ — $50+((68-22)/20) \cdot 5=61.5$ uses the full $n=68$ instead of $n/2=34$, giving a far larger, wrong figure.
$48$ kg, from using the running total through 50-55 itself, $C=42$, instead of the total before it — $C$ must be the total BEFORE the median class, $22$ — using $42$, the total through it, gives $50+((34-42)/20) \cdot 5=48$, the wrong value.
$50$ kg, the lower limit of the median class, mistaken for the median — $50$ is just the lower limit, $l$ — the median formula shifts further into the class, to $53$.
Why does the median formula use $C$, the cumulative frequency up to but not including the median class, rather than the cumulative total through the median class?
either cumulative total gives the same answer, since $f$ absorbs the difference
$(n/2-C)$ measures how far into the median class the middle value sits
$C$ is defined that way only to keep the formula’s units consistent with $h$
the cumulative total through the median class is not available once the median class itself is known
Check your answer
either cumulative total gives the same answer, since $f$ absorbs the difference — Using the cumulative total through the median class instead of before it changes $(n/2-C)$, and so changes the computed median — the two are not interchangeable.
✓ $(n/2-C)$ measures how far into the median class the middle value sits — (B) $(n/2-C)$ is the number of observations still needed once the classes before the median class are counted — that is exactly how far into the median class the middle value sits.
$C$ is defined that way only to keep the formula’s units consistent with $h$ — $C$ and $h$ have nothing to do with matching units — $C$ is defined to measure frequency accumulated strictly BEFORE the median class.
the cumulative total through the median class is not available once the median class itself is known — Both cumulative totals — before, and through, the median class — are available directly from the table; the formula simply needs the one measured before it.
50 plant heights are grouped 0-20, 20-40, 40-60, 60-80, 80-100 with frequencies 5, 8, 12, 15, 10. Here $n/2=25$, the median class is 40-60 with $l=40$, $h=20$, $C=13$, $f=12$. The median is
$60$
$101.67$, from using $n=50$ instead of $n/2=25$
$40$, from using the running total through 40-60 itself, $C=25$, instead of the total before it
$50$, the class mark of the median class, mistaken for the median
$101.67$, from using $n=50$ instead of $n/2=25$ — $40+((50-13)/12) \cdot 20=101.67$ uses the full $n=50$ instead of $n/2=25$, giving a value far outside the median class itself.
$40$, from using the running total through 40-60 itself, $C=25$, instead of the total before it — $C$ must be the total BEFORE the median class, $13$ — using $25$, the total through it, gives $40+((25-25)/12) \cdot 20=40$, the wrong value.
$50$, the class mark of the median class, mistaken for the median — $50$ is the midpoint of 40-60 — the median formula shifts further into the class, to $60$.
The median of grouped data comes from one formula, applied to the median class. It is $l + ((n/2 - C)/f) \cdot h$. Five numbers go into this formula, and four of them are already on the frequency table.
Here $l$ is the median class’s lower limit and $n$ is the total number of observations. $C$ is the running total of frequencies before that class. $f$ is the median class’s own frequency, and $h$ is the class size. We build the running total, class by class, and stop the moment it first reaches $n/2$.
Get $C$ wrong and every later number in the solution shifts too.
This formula assumes every class has the same size, $h$, throughout the table. We check our own table for this before trusting the formula on it.
Your turn: a median class has $l = 15$, $n = 50$, $C = 18$, $f = 10$ and $h = 5$. What is the median? (Answer: $18.5$, since $n/2 = 25$ and $15 + ((25 - 18)/10) \cdot 5 = 15 + 3.5 = 18.5$.)
In the median class 30 to 40, the formula’s answer, 34.67, sits close to but not exactly at the class midpoint, 35.
A student computes the mean of homework-minutes data (classes 0-10 to 40-50, frequencies 3, 5, 8, 3, 1) twice by the assumed-mean method: once with $a=25$, getting $\overline{x}=22$, and once with $a=15$. The student expects a different answer the second time, since a different $a$ was chosen. What actually happens?
the second calculation gives a different mean, since $a=15$ is a smaller number than $a=25$
the second calculation also gives $\overline{x}=22$ — any valid choice of $a$ gives the same mean
the second calculation cannot be completed, since $a=15$ does not appear as a class mark in this table
the second calculation gives $\overline{x}=32$, since the smaller $a$ needs $10$ added back to match the first result
Check your answer
the second calculation gives a different mean, since $a=15$ is a smaller number than $a=25$ — Every valid choice of $a$ leads back to the same $\overline{x}=22$ here — changing $a$ only changes the intermediate deviations $d_i$, not the final mean.
✓ the second calculation also gives $\overline{x}=22$ — any valid choice of $a$ gives the same mean — (B) Choosing $a=15$ instead of $a=25$ still gives $\overline{x}=22$ — any valid choice of the assumed mean leads back to the same result.
the second calculation cannot be completed, since $a=15$ does not appear as a class mark in this table — $a$ is commonly chosen as a class mark for convenience, but the method works for any value of $a$, not only ones that appear in the table.
the second calculation gives $\overline{x}=32$, since the smaller $a$ needs $10$ added back to match the first result — There is no adjustment to add between choices of $a$ — the assumed-mean formula already accounts for whichever $a$ is chosen, and both give exactly $22$.
Why does changing the assumed mean $a$ leave the final $\overline{x}$ unchanged, even though every deviation $d_i=x_i-a$ changes?
because $\sum f_i$ also changes to compensate, keeping the ratio fixed
because the class marks $x_i$ change along with $a$
the shift in $d_i$ from a new $a$ cancels when $a$ is added back
because only a very small range of values of $a$ actually give the correct mean
Check your answer
because $\sum f_i$ also changes to compensate, keeping the ratio fixed — $\sum f_i$ is the total number of observations — it is fixed by the data and does not change with the choice of $a$.
because the class marks $x_i$ change along with $a$ — The class marks $x_i$ come from the table’s own class limits and never change — only the deviations $d_i=x_i-a$ shift when $a$ changes.
✓ the shift in $d_i$ from a new $a$ cancels when $a$ is added back — (C) A different $a$ shifts every deviation by the same amount, and adding that same $a$ back at the end undoes exactly that shift.
because only a very small range of values of $a$ actually give the correct mean — Any value of $a$ gives the same correct mean once the correction term is added back — there is no narrow range it must fall within.
A different grouped table gives $\overline{x}=64$ when computed by the assumed-mean method with $a=60$. Without recomputing, what mean should the same data give if reworked with $a=50$ instead?
$54$, since the mean should shift down by the same $10$ that $a$ dropped by
it cannot be known without redoing the full calculation
$74$, since the mean should shift up by $10$ to compensate for the smaller $a$
$64$ — the same mean, whatever valid $a$ is chosen
Check your answer
$54$, since the mean should shift down by the same $10$ that $a$ dropped by — The mean of the underlying data does not shift with a new choice of $a$ — both calculations correct back to the same value, $64$.
it cannot be known without redoing the full calculation — That the mean stays fixed under any valid $a$ is exactly what lets it be known here without redoing the arithmetic — it is still $64$.
$74$, since the mean should shift up by $10$ to compensate for the smaller $a$ — The mean of the underlying data does not shift with a new choice of $a$ — both calculations correct back to the same value, $64$.
✓ $64$ — the same mean, whatever valid $a$ is chosen — (D) The mean of a fixed dataset does not depend on which assumed mean was used to compute it — reworking with $a=50$ still gives $64$.
TRAP: choosing a different value for the assumed mean, $a$, changes the mean.
REALITY: it does not. Any valid choice of $a$ gives exactly the same mean.
*If we move $a$ up by any amount, every deviation moves down by that same amount.* The two changes cancel, so the mean stays the same. We pick $a$ near the middle of the class marks only to keep the deviations small, never because one value is more correct than another.
Your turn: the homework table (frequencies $3, 5, 8, 3, 1$, marks $5$ to $45$) takes $a = 35$. What mean does it give? (Answer: $22$, since the deviations are $-30, -20, -10, 0, 10$, their weighted total is $-260$, and $35 + (-260)/20 = 35 - 13 = 22$, the same $22$ as every other route.)
Does changing the assumed mean change the mean you compute?
Weaker. Two readers work the same table: classes 0–10, 10–20, 20–30, 30–40 with frequencies $2, 3, 4, 1$, so $n = 10$ and class marks $5, 15, 25, 35$. The first reader takes $a = 15$: the deviations are $-10, 0, 10, 20$, their weighted total is $40$, and the mean is $15 + 40/10 = 19$. The second reader takes $a = 25$, but keeps the deviations already worked out from $15$ instead of recalculating them, so gets $25 + 40/10 = 29$, and decides the choice of $a$ changed the mean.
A correct working cannot depend on $a$, so two different means from two values of $a$ mean one working is wrong. Moving $a$ up by $10$ must move every deviation down by $10$.
Stronger. With $a = 25$ the deviations are $-20, -10, 0, 10$, their weighted total is $-60$, and the mean is $25 + (-60)/10 = 19$.
Both values of $a$ give $19$, the same as the direct method’s $190/10 = 19$.
A SCHOLAR INDIA REMEMBERS
Do you need the perfect assumed mean? Test it. Work the mean once with $30$ as the assumed mean, and once with $40$. You get the same mean both times. So the choice only changes how big the numbers are while you work. A class mark near the middle keeps them small.
Read the frequencies either side of the modal class from the right rows
✕MISCONCEPTION
Check yourself
For TV-watching hours grouped 0-5, 5-10, 10-15, 15-20, 20-25 with frequencies 2, 6, 10, 5, 2, the modal class is 10-15 with $f_1=10$. A student sets $f_0=10$, reading it off the modal class’s own row, instead of $6$, the frequency of the class immediately before it. What mode does each approach give?
both methods give the same mode, $12.22$, since $f_0$ cancels out of the formula either way
the correct method gives mode $=10$; the student’s method gives mode $=12.22$ instead
the correct method gives mode $=12.22$; the student’s method gives mode $=10$ instead
the correct method gives mode $=12.22$; the student’s method gives mode $=9$ instead
Check your answer
both methods give the same mode, $12.22$, since $f_0$ cancels out of the formula either way — $f_0$ appears directly in the formula’s numerator and denominator — using the wrong row for it changes the computed mode, from $12.22$ down to $10$.
the correct method gives mode $=10$; the student’s method gives mode $=12.22$ instead — It is the correct reading, $f_0=6$, that gives $12.22$ — reusing the modal row’s own frequency as $f_0=10$ is what gives the wrong value, $10$.
✓ the correct method gives mode $=12.22$; the student’s method gives mode $=10$ instead — (C) With the correct $f_0=6$: $10+((10-6)/(20-6-5)) \cdot 5=12.22$. With the student’s $f_0=10$: $10+((10-10)/(20-10-5)) \cdot 5=10$.
the correct method gives mode $=12.22$; the student’s method gives mode $=9$ instead — Setting $f_0=10$ actually gives $10+((10-10)/(20-10-5)) \cdot 5=10$, not $9$ — recompute the student’s own numbers rather than guessing.
In the mode formula, why must $f_0$ come from the class immediately before the modal class, rather than any other class in the table?
it compares the modal class against its two immediate neighbours only
because $f_0$ must always be the smallest frequency in the entire table
because the classes must be renumbered so the modal class becomes the first row
because $f_0$ and $f_1$ must always be equal for the formula to apply
Check your answer
✓ it compares the modal class against its two immediate neighbours only — (A) The formula is built entirely around the modal class and its two direct neighbours — no other class in the table plays any role in it.
because $f_0$ must always be the smallest frequency in the entire table — $f_0$ is simply the frequency of the class immediately before the modal one — it need not be the smallest frequency anywhere in the table.
because the classes must be renumbered so the modal class becomes the first row — The classes keep their original order in the table — $f_0$ and $f_2$ are simply read off whichever rows sit directly before and after the modal class.
because $f_0$ and $f_1$ must always be equal for the formula to apply — $f_0$ and $f_1$ are two different classes’ frequencies and need not be equal — treating them as equal is exactly the slip that gives a wrong mode.
Marks are grouped 0-10, 10-20, 20-30, 30-40, 40-50 with frequencies 4, 12, 20, 9, 5. The modal class is 20-30. Which frequencies are the correct $f_0$ and $f_2$ for the mode formula?
$f_0=20$ and $f_2=20$, both read off the modal class’s own row
$f_0=9$ and $f_2=12$, with the two neighbours swapped
$f_0=12$ (the class before, 10-20) and $f_2=9$ (the class after, 30-40)
$f_0=4$ and $f_2=5$, the frequencies of the first and last classes in the table
Check your answer
$f_0=20$ and $f_2=20$, both read off the modal class’s own row — $f_0$ and $f_2$ must come from the two classes bordering the modal class, 10-20 and 30-40 — not from the modal class’s own row.
$f_0=9$ and $f_2=12$, with the two neighbours swapped — $f_0$ is the class BEFORE the modal class ($10$-$20$, frequency $12$), and $f_2$ the class AFTER it ($30$-$40$, frequency $9$) — swapped, these give a different, wrong mode.
✓ $f_0=12$ (the class before, 10-20) and $f_2=9$ (the class after, 30-40) — (C) $f_0$ is the frequency of the class directly before the modal class, and $f_2$ the frequency of the class directly after it: $12$ and $9$.
$f_0=4$ and $f_2=5$, the frequencies of the first and last classes in the table — The first and last classes in the table, 0-10 and 40-50, are not adjacent to the modal class, 20-30 — only its immediate neighbours, 10-20 and 30-40, supply $f_0$ and $f_2$.
TRAP: read $f_0$ off the modal class’s own row, instead of the row immediately before it.
REALITY: that gives a wrong mode. $f_0$ comes from the row just above the modal class, never from the modal row itself.
*If we reuse the modal row’s own frequency as $f_0$, the top of the fraction becomes zero.* The mode then lands exactly on the lower limit $l$, not inside the modal class. Copying $f_1$ into $f_0$ or $f_2$ out of habit is the single most common slip in this formula.
Your turn: a modal class 40–50 has $f_1 = 14$, $f_0 = 8$ (the row before) and $f_2 = 10$ (the row after), with $h = 10$. What is the mode? (Answer: $46$, since $40 + ((14 - 8)/(28 - 8 - 10)) \cdot 10 = 40 + 6 = 46$.)
Reusing the modal row’s own frequency as f0 drags the mode from 46 down onto the class’s lower limit, 40.
First try
Modal class $10$ to $15$, with $f_1 = 10$ and $f_2 = 5$. I read $f_0$ straight off the same row, $f_0 = 10$, and got mode $= 10 + (10 - 10)/(20 - 10 - 5) \cdot 5 = 10$.
Second look
$f_0$ comes from the class just before the modal class, not the modal class itself. That class had frequency $6$, so $f_0 = 6$, and mode $= 10 + (10 - 6)/(20 - 6 - 5) \cdot 5 = 10 + 4/9 \cdot 5 \approx 12.22$.
$f_0$ always comes from the row just above the modal class, never from the modal row.
Which two frequencies feed the mode formula, the modal row or its neighbours?
Weaker. For classes 0–10, 10–20, 20–30, 30–40, 40–50 with frequencies $3, 7, 12, 8, 2$, the modal class is 20–30, since $12$ is the highest frequency, so $l = 20$, $h = 10$, $f_1 = 12$. A reader takes $f_0 = 12$ as well, reading it from the modal row itself instead of the row before it. That makes the top of the fraction $12 - 12 = 0$, so the mode comes out as $20 + 0 = 20$, exactly on the lower limit.
The row before the modal class holds $7$, not $12$, so $f_0 = 12$ cannot be right. A mode landing exactly on the modal class’s lower limit, with nothing pulling it inward, is a sign the wrong frequency went in.
Stronger. $f_0$ is the frequency of the class before the modal class, and $f_2$ is the frequency of the class after it. Here $f_0 = 7$ and $f_2 = 8$.
$20 + ((12 - 7)/(24 - 7 - 8)) \cdot 10 \approx 25.56$, a value inside the modal class, between $20$ and $30$, as it must be.
20 students’ homework time is grouped 0-10, 10-20, 20-30, 30-40, 40-50 minutes with frequencies 3, 5, 8, 3, 1. By the direct method, $\sum f_i x_i =$
$20$, the value of $\sum f_i$, mistaken for $\sum f_i x_i$
$125$, the sum of the class marks alone, without weighting by frequency
$22$, the final mean, mistaken for the intermediate sum
$440$
Check your answer
$20$, the value of $\sum f_i$, mistaken for $\sum f_i x_i$ — $20$ is $\sum f_i$, the count of students — the weighted sum $\sum f_i x_i$, needed for the mean, comes to $440$.
$125$, the sum of the class marks alone, without weighting by frequency — $5+15+25+35+45=125$ ignores the frequencies entirely — weighting each mark by its own frequency gives $440$.
$22$, the final mean, mistaken for the intermediate sum — $22$ is the FINAL mean, $440/20$ — the intermediate weighted sum on the way there is $440$.
In this worked example, why does the class 20-30 contribute the largest single term, $8(25)=200$, to the running total $\sum f_i x_i$?
20-30 has the highest frequency, 8, so it is weighted most
because 20-30 is the middle class in the table, and the middle class always contributes the most
because $25$ is the largest class mark among the five classes
because the direct method always assigns extra weight to whichever class the assumed mean is chosen from
Check your answer
✓ 20-30 has the highest frequency, 8, so it is weighted most — (A) 20-30 has the highest frequency of the five classes, so its mark, $25$, gets multiplied by the largest weight, giving the largest single term.
because 20-30 is the middle class in the table, and the middle class always contributes the most — Table position decides nothing here — 20-30 contributes the most because it has the highest FREQUENCY, $8$, not because it sits in the middle.
because $25$ is the largest class mark among the five classes — $45$, not $25$, is the largest class mark here — but its frequency is only $1$, giving it a small term; frequency, not mark size, decides the largest contribution.
because the direct method always assigns extra weight to whichever class the assumed mean is chosen from — The direct method never chooses an assumed mean at all — every class is weighted purely by its own frequency.
A second group of 25 students records $\sum f_i x_i = 500$ over the same five class intervals as the homework data ($\sum f_i x_i=440$, $\sum f_i=20$ for the first group). Which group has the higher mean homework time?
the second group, since $500>440$
both groups have the same mean, since they cover the same five class intervals
the first group, since $440/20=22$ exceeds $500/25=20$
the second group, since it has more students overall
Check your answer
the second group, since $500>440$ — Comparing the totals directly ignores that the groups have different sizes — dividing each by its own $\sum f_i$ shows the first group’s mean, $22$, is actually higher.
both groups have the same mean, since they cover the same five class intervals — Sharing the same class intervals says nothing about how the students are distributed across them — the two groups’ means, $22$ and $20$, differ.
✓ the first group, since $440/20=22$ exceeds $500/25=20$ — (C) Dividing each group’s weighted sum by its OWN total frequency: $440/20=22$ for the first group, $500/25=20$ for the second — the first group’s mean is higher, despite its smaller total.
the second group, since it has more students overall — Having more students says nothing about the mean homework time — dividing each group’s total by its own size gives $22$ for the first group and $20$ for the second.
Worked example
Mean of homework time by the direct method
classes 0–10, 10–20, 20–30, 30–40, 40–50 minutes; frequencies $3, 5, 8, 3, 1$; class marks $5, 15, 25, 35, 45$ class marks read as the midpoint of each interval, from the frequency table for 20 students’ homework time
$\sum f_i x_i = 3(5) + 5(15) + 8(25) + 3(35) + 1(45) = 15 + 75 + 200 + 105 + 45 = 440$ multiply each class mark by its own frequency, then add every product
$\sum f_i = 20$ add the frequency column to get the total number of students
$\overline{x} = 440/20 = 22$ minutes divide the total of $f_i x_i$ by the total frequency, the direct-method formula
check by the assumed-mean method with $a = 25$: $\sum f_i d_i = 3(-20) + 5(-10) + 8(0) + 3(10) + 1(20) = -60 - 50 + 0 + 30 + 20 = -60$ Multiply each class mark’s deviation from $a = 25$ by its own frequency, then add every product: the same homework-time table, a different route to the mean.
$\overline{x} = 25 + (-60/20) = 22$ minutes Check the direct-method answer by a second route. The assumed-mean method lands on the same $22$ minutes, so the total added in step 2 was right.
The mean is the one point on the ruler where the classes above and the classes below pull with equal force.
Find the mean of grouped data by the direct method.
List the class marks Write down each class’s class mark. For classes 0–10 to 40–50, the marks are $5, 15, 25, 35, 45$.
Multiply by frequency The frequencies are $3, 5, 8, 3, 1$. Multiply each class mark by its own frequency, giving $5(3) = 15$, $15(5) = 75$, $25(8) = 200$, $35(3) = 105$, $45(1) = 45$.
Add both columns Add the frequencies to get the total count, $20$. Add the five products to get $440$.
Divide Divide the total of the products by the total frequency. $440/20 = 22$ minutes.
Check the size The mean must sit among the class marks, closest to the tallest class. $22$ sits between $15$ and $25$, near the class 20–30, the one with the highest frequency.
50 tea packets are grouped 100-120, 120-140, 140-160, 160-180, 180-200 g with frequencies 4, 10, 20, 12, 4. Taking $a=150$, the deviations are $d_i=-40,-20,0,20,40$, and $\sum f_i d_i=40$. The mean weight is
$150$ g, the assumed mean with no correction added
$0.8$ g, the correction term alone, without adding back $a$
$190$ g, adding $a$ and $\sum f_i d_i$ with the division by $\sum f_i$ skipped
$150.8$ g
Check your answer
$150$ g, the assumed mean with no correction added — $150$ g is only the assumed mean — the correction, $0.8$ g, still needs to be added to reach $150.8$ g.
$0.8$ g, the correction term alone, without adding back $a$ — $0.8$ g is just $40/50$ — adding it back to $a=150$ g gives the actual mean, $150.8$ g.
$190$ g, adding $a$ and $\sum f_i d_i$ with the division by $\sum f_i$ skipped — $40$ must be divided by $\sum f_i=50$ before adding it to $a$ — adding it raw gives $150+40=190$ g, far too large.
✓ $150.8$ g — (D) $\overline{x} = 150 + 40/50 = 150+0.8=150.8$ g.
Which single number in this worked example plays the role of $a$, the assumed mean, chosen for convenience rather than computed?
$40$, the frequency-weighted deviation total $\sum f_i d_i$
$0.8$, the correction added back to $a$
$150.8$, the final computed mean
$150$, the class mark of the middle class, 140-160
Check your answer
$40$, the frequency-weighted deviation total $\sum f_i d_i$ — $40$ is $\sum f_i d_i$, a quantity computed AFTER choosing $a$ — it is not $a$ itself.
$0.8$, the correction added back to $a$ — $0.8$ is the correction, $\sum f_i d_i/\sum f_i$ — it is added TO $a$, not the same number as $a$.
$150.8$, the final computed mean — $150.8$ is the FINAL mean, the method’s output — $a=150$ is the starting choice the method corrects.
✓ $150$, the class mark of the middle class, 140-160 — (D) $a=150$ is simply picked as a convenient class mark, here the middle one — it is not itself computed from the data.
Using this worked example’s numbers, suppose a slip records $\sum f_i d_i$ as $-40$ instead of the correct $40$ (a sign error). What mean would that give?
$150.8$ g, the same as the correct mean, since the sign of $\sum f_i d_i$ does not affect the final answer
$-149.2$ g, treating the whole mean as negative
$110$ g, subtracting $40$ from $a$ with no division by $\sum f_i$ at all
$150 + (-40)/50 = 149.2$ g
Check your answer
$150.8$ g, the same as the correct mean, since the sign of $\sum f_i d_i$ does not affect the final answer — The sign of $\sum f_i d_i$ directly flips the sign of the correction term — $150+0.8=150.8$ g with the correct sign, but $150-0.8=149.2$ g with the flipped one.
$-149.2$ g, treating the whole mean as negative — Only the CORRECTION term flips sign, not the whole mean — $a=150$ g stays positive, giving $150-0.8=149.2$ g, not a negative mean.
$110$ g, subtracting $40$ from $a$ with no division by $\sum f_i$ at all — $-40$ must be divided by $\sum f_i=50$ before adding it to $a$ — subtracting it raw gives $150-40=110$ g, far too large a shift.
✓ $150 + (-40)/50 = 149.2$ g — (D) Flipping the sign of $\sum f_i d_i$ flips the sign of the correction term too: $150+(-0.8)=149.2$ g.
A stacked column of five weighted blocks totals 7540 grams, giving a mean of 150.8 grams over 50 packets.
Worked example
Mean tea-packet weight by the assumed-mean method
classes 100–120 to 180–200 grams; frequencies $4, 10, 20, 12, 4$; class marks $110, 130, 150, 170, 190$ class marks read from the frequency table for 50 tea packets
take $a = 150$; deviations $d_i = -40, -20, 0, 20, 40$ choose the middle class mark as the assumed mean, then subtract it from every class mark
$\sum f_i d_i = 4(-40) + 10(-20) + 20(0) + 12(20) + 4(40) = -160 - 200 + 0 + 240 + 160 = 40$ multiply each deviation by its own frequency, then add every product
$\overline{x} = 150 + 40/50 = 150.8$ grams add the assumed mean to the total deviation divided by the total frequency
check by the direct method: $\sum f_i x_i = 4(110) + 10(130) + 20(150) + 12(170) + 4(190) = 7540$ multiply each class mark by its own frequency directly, without any deviation shortcut
$7540/50 = 150.8$ grams Check the assumed-mean answer against the direct method. Both land on $150.8$ grams, so the shortcut changed nothing.
The two classes below the assumed mean nearly cancel the two above, and the small orange bar is what moves the answer off 150 grams.
Find the mean of grouped data by the assumed-mean method.
List the class marks Write down each class’s class mark. For the tea-packet classes, the marks are $110, 130, 150, 170, 190$.
Choose the assumed mean Pick a class mark near the middle as $a$. Here $a = 150$.
Find each deviation Subtract $a$ from every class mark. $d_i = -40, -20, 0, 20, 40$.
Multiply and add The frequencies are $4, 10, 20, 12, 4$. Multiply each deviation by its own frequency and add the products. $4(-40) + 10(-20) + 20(0) + 12(20) + 4(40) = 40$.
Add back to a Add the frequencies to get $50$. Divide the deviation total by $50$ and add it to $a$. $150 + 40/50 = 150.8$ grams.
Check with a second route Work the same table with a different assumed mean, or by the direct method. Either route must give back the same $150.8$ grams.
60 delivery riders’ daily distance is grouped 0-20, 20-40, 40-60, 60-80, 80-100 km with frequencies 6, 14, 22, 12, 6. Taking $a=50$, $h=20$ gives $u_i=-2,-1,0,1,2$ and $\sum f_i u_i=-2$. The mean distance is
$49.33$ km
$50$ km, the plain assumed mean value
$49.97$ km, from leaving out the multiplication by $h$
$10$ km, adding $a$ to $h$ times the raw total, division by $\sum f_i$ omitted
Check your answer
✓ $49.33$ km — (A) $\overline{x} = 50 + 20 \cdot (-2/60) = 50-0.67=49.33$ km.
$50$ km, the plain assumed mean value — $50$ km is only the assumed mean — the correction, $-0.67$ km, still needs to be added to reach $49.33$ km.
$49.97$ km, from leaving out the multiplication by $h$ — $50+(-2/60)=49.97$ km leaves out multiplying the correction by $h=20$ — doing so gives the true mean, $49.33$ km.
$10$ km, adding $a$ to $h$ times the raw total, division by $\sum f_i$ omitted — $20 \cdot (-2)=-40$ must be divided by $\sum f_i=60$ before adding it to $a$ — adding it raw gives $50-40=10$ km, far too large a shift.
In this worked example, $\sum f_i u_i=-2$ comes out negative. What does that indicate about the mean, before finishing the calculation?
the calculation has an error, since $\sum f_i u_i$ can never be negative
the mean will land slightly below $a=50$
the mean will be far below $a$, by roughly $h=20$ km
the mean will be above $a=50$, since negative values always increase a mean
Check your answer
the calculation has an error, since $\sum f_i u_i$ can never be negative — $\sum f_i u_i$ can be negative whenever more of the data sits below $a$ than above it — it is not a sign of an error here.
✓ the mean will land slightly below $a=50$ — (B) A negative $\sum f_i u_i$ means the correction term, once divided by $\sum f_i$ and multiplied by $h$, comes out negative — so the mean lands just below $a$.
the mean will be far below $a$, by roughly $h=20$ km — The shift is $h \cdot \sum f_i u_i/\sum f_i = 20 \cdot (-2)/60$, only about $0.67$ km — nowhere near a full class width of $20$ km.
the mean will be above $a=50$, since negative values always increase a mean — Adding a NEGATIVE correction term to $a$ moves the mean DOWN, not up — the mean here comes out below $50$, at $49.33$.
A second group of 40 riders gives $\sum f_i u_i = 6$ with the same $a=50$, $h=20$. What mean distance does this give?
$50$ km, the assumed mean, correction never added
$50.15$ km, from leaving out the multiplication by $h$
$170$ km, adding $a$ to $h$ times the raw total, with $\sum f_i$ never applied
$53$ km
Check your answer
$50$ km, the assumed mean, correction never added — $50$ km is only the assumed mean — the correction, $3$ km, still needs to be added to reach $53$ km.
$50.15$ km, from leaving out the multiplication by $h$ — $50+6/40=50.15$ km leaves out multiplying the correction by $h=20$ — doing so gives the true mean, $53$ km.
$170$ km, adding $a$ to $h$ times the raw total, with $\sum f_i$ never applied — $20 \cdot 6=120$ must be divided by $\sum f_i=40$ before adding it to $a$ — adding it raw gives $50+120=170$ km, far too large.
✓ $53$ km — (D) $\overline{x} = 50 + 20 \cdot (6/40) = 50+3=53$ km.
Worked example
Mean delivery distance by the step-deviation method
classes 0–20 to 80–100 km; frequencies $6, 14, 22, 12, 6$; class marks $10, 30, 50, 70, 90$ class marks read from the frequency table for 60 delivery riders
take $a = 50$ and $h = 20$; $u_i = -2, -1, 0, 1, 2$ every deviation from $a = 50$ shares the common factor $h = 20$, so divide it out
$\sum f_i u_i = 6(-2) + 14(-1) + 22(0) + 12(1) + 6(2) = -12 - 14 + 0 + 12 + 12 = -2$ multiply each $u_i$ by its own frequency, then add every product
$\overline{x} = 50 + 20 \cdot (-2/60) \approx 49.33$ km the step-deviation formula, multiplying back by $h$ before adding to $a$
check by the direct method: $\sum f_i x_i = 6(10) + 14(30) + 22(50) + 12(70) + 6(90) = 2960$ multiply each class mark by its own frequency directly, without dividing out the class size
$2960/60 \approx 49.33$ km Check the step-deviation answer against the direct method. Both land on $49.33$ km: it is a shortcut, not a different answer.
The same five bars carry two rows of numbers, class widths and kilometres, both giving a mean of 49.33 km.
Find the mean of grouped data by the step-deviation method.
List the class marks Write down each class’s class mark. For the delivery classes, the marks are $10, 30, 50, 70, 90$.
Choose a and read h Pick a class mark near the middle as $a$, and read the class size $h$ from the table. Here $a = 50$ and $h = 20$.
Divide each deviation by h Subtract $a$ from each class mark, then divide by $h$. $u_i = -2, -1, 0, 1, 2$.
Multiply and add The frequencies are $6, 14, 22, 12, 6$. Multiply each $u_i$ by its own frequency and add the products. $6(-2) + 14(-1) + 22(0) + 12(1) + 6(2) = -2$.
Scale back up and add to a Add the frequencies to get $60$. Multiply the total from the step before by $h$, divide by $60$, and add to $a$. $50 + 20 \cdot (-2/60) \approx 49.33$ km.
Check by the direct method Work the same table by the direct method. It must give back the same $\approx 49.33$ km.
A missing frequency, found from a mean you already know
classes 10–20 to 50–60 pots; frequencies $2, k, 13, 5, 2$; class marks $15, 25, 35, 45, 55$ the clay pots fired in a day by a group of potters, with one frequency never recorded
$\sum f_i x_i = 2(15) + 13(35) + 5(45) + 2(55) + 25k = 820 + 25k$ every class you can count goes in as a number, and the one you cannot keeps its letter
$\sum f_i = 2 + 13 + 5 + 2 + k = 22 + k$ the total is short by the same unknown, so it carries the letter too
$820 + 25k = 34(22 + k)$ the mean is given as $34$ pots, and the mean times the total equals the sum of $f_i x_i$
$820 + 25k = 748 + 34k$, so $72 = 9k$ and $k = 8$ multiply out the bracket, then gather the letters on one side and the numbers on the other
check by putting $k = 8$ back. $\sum f_i x_i = 820 + 200 = 1020$, $\sum f_i = 30$, and $1020/30 = 34$ *Check the mean comes back as the $34$ pots the question gave.* If it did not, the arithmetic went wrong, not the method.
one row
the pot she is counting
Ananya counts the pots drying in a potter’s courtyard, row by row.
Find a missing frequency from a mean you already know.
Write the unknown into the table Give the missing frequency a letter, $k$. For the pots table, the frequencies are $2, k, 13, 5, 2$ with class marks $15, 25, 35, 45, 55$.
Total the known frequencies Add the frequencies that are known numbers. $2 + 13 + 5 + 2 = 22$, so the total frequency is $22 + k$.
Total the known products Multiply the known frequencies by their class marks and add. $2(15) + 13(35) + 5(45) + 2(55) = 820$, plus $25k$ from the unknown class.
Set mean times total equal to the sum The mean is $34$ pots. Write $820 + 25k = 34(22 + k)$.
Solve for k Multiply out the bracket and collect terms. $820 + 25k = 748 + 34k$, so $72 = 9k$ and $k = 8$.
Check by putting k back Put $k = 8$ back into the total frequency and the total of the products, then divide. The mean must come back as $34$.
Two equal-length bars for 820 plus 25k and 748 plus 34k close a gap of 72, giving k equals 8.
25 people’s weekly TV hours are grouped 0-5, 5-10, 10-15, 15-20, 20-25 with frequencies 2, 6, 10, 5, 2. The modal class is 10-15, since
its class mark, $12.5$, is closest to the mean of all the class marks
its frequency, $10$, is the highest in the table
it has the lowest frequency among its two neighbouring classes
it is the only class with a frequency greater than 5
Check your answer
its class mark, $12.5$, is closest to the mean of all the class marks — The modal class is chosen purely by comparing frequencies — proximity of its class mark to any mean is not part of the rule.
✓ its frequency, $10$, is the highest in the table — (B) 10-15 carries frequency $10$, higher than any other class in the table, so it is the modal class.
it has the lowest frequency among its two neighbouring classes — 10-15 has the HIGHEST frequency in the whole table, $10$ — its neighbours, at $6$ and $5$, are both lower, not higher.
it is the only class with a frequency greater than 5 — 5-10 also has frequency $6$, greater than $5$ — being the only such class is not why 10-15 is modal; having the single highest frequency, $10$, is.
For the TV-hours data, the modal class 10-15 has $l=10$, $h=5$, $f_1=10$, $f_0=6$, $f_2=5$. The mode is
$10$ hours, from wrongly reusing $f_0=10$ instead of $6$
$12.78$ hours, from swapping $f_0$ and $f_2$ in the formula
$12.22$ hours
$12.5$ hours, the class mark of the modal class, mistaken for the mode
Check your answer
$10$ hours, from wrongly reusing $f_0=10$ instead of $6$ — Reusing $f_0=10$ (the modal class’s own frequency) gives $10+((10-10)/(20-10-5)) \cdot 5=10$ hours — the correct $f_0=6$ gives $12.22$ hours.
$12.78$ hours, from swapping $f_0$ and $f_2$ in the formula — Swapping the neighbours gives $10+((10-5)/(20-5-6)) \cdot 5=12.78$ hours — the correct reading, $f_0=6$ and $f_2=5$, gives $12.22$ hours.
$12.5$ hours, the class mark of the modal class, mistaken for the mode — $12.5$ is the midpoint of 10-15 — the mode formula shifts away from that midpoint toward the nearer-frequency neighbour, landing on $12.22$ hours.
Suppose the frequency of the class before the modal class (5-10) had been recorded as $8$ instead of $6$, with everything else the same ($l=10$, $h=5$, $f_1=10$, $f_2=5$). What would the mode become?
$12.22$ hours, unchanged from before, since only $f_0$ was edited
$11.11$ hours, from keeping the old denominator ($9$) instead of recomputing it with the new $f_0=8$
$11.43$ hours
$14.22$ hours, from adding the frequency difference ($8-6=2$) directly onto the old mode
Check your answer
$12.22$ hours, unchanged from before, since only $f_0$ was edited — $f_0$ appears in both the numerator and the denominator of the mode formula — changing it from $6$ to $8$ does change the mode, to $11.43$ hours.
$11.11$ hours, from keeping the old denominator ($9$) instead of recomputing it with the new $f_0=8$ — The denominator, $2f_1-f_0-f_2$, must also use the new $f_0=8$, giving $7$, not the old $9$ — using $9$ gives $11.11$ hours instead of the correct $11.43$.
✓ $11.43$ hours — (C) $10 + ((10-8)/(20-8-5)) \cdot 5 = 10+(2/7) \cdot 5 = 11.43$ hours — both the numerator and the denominator must use the new $f_0=8$.
$14.22$ hours, from adding the frequency difference ($8-6=2$) directly onto the old mode — There is no shortcut that adds the change in $f_0$ straight onto the old mode — the formula must be recomputed in full, giving $11.43$ hours.
Worked example
Mode of weekly TV-watching hours
classes 0–5 to 20–25 hours; frequencies $2, 6, 10, 5, 2$ the frequency table for 25 people’s weekly TV-watching hours
highest frequency $10$ falls in the class 10–15 — the modal class scan the frequency column and pick its largest entry
$l = 10$, $h = 5$, $f_1 = 10$, $f_0 = 6$, $f_2 = 5$ read the modal class’s lower limit and class size, and its own and neighbouring frequencies, straight off the table
mode $= 10 + ((10 - 6)/(20 - 6 - 5)) \cdot 5 = 10 + (4/9) \cdot 5 \approx 12.22$ hours substitute into the mode formula for grouped data
$10 < 12.22 < 15$ Check the size of the answer. The mode must sit inside the modal class it came from, never outside it, and here it does.
The two step-ups into the modal class split its base in their own ratio, so the mode leans toward the taller neighbour.
Find the mode of grouped data.
Find the modal class Scan the frequency column for the highest count. For the TV-watching table, frequencies $2, 6, 10, 5, 2$ peak at $10$, in the class 10–15.
Read l, h and f one Read the modal class’s lower limit, class size and own frequency. $l = 10$, $h = 5$, $f_1 = 10$.
Read f nought and f two Read the frequencies of the row before and the row after the modal class. $f_0 = 6$, $f_2 = 5$.
Substitute Put the five numbers into the mode formula. $10 + ((10 - 6)/(20 - 6 - 5)) \cdot 5 \approx 12.22$ hours.
Check the size The mode must sit inside the modal class, between $10$ and $15$. $12.22$ does.
40 fitness-club members’ ages are grouped 10-20, 20-30, 30-40, 40-50, 50-60 with frequencies 4, 9, 15, 8, 4. Here $n/2=20$, and the median class is 30-40, since
it is the class where the running total first EQUALS exactly $20$
the running total first passes $20$ there, at $28$
it is the last class where the running total is below $20$
it is determined by comparing class marks to the mean, not by a running total
Check your answer
it is the class where the running total first EQUALS exactly $20$ — The running total jumps from $13$ to $28$ across this class, never landing exactly on $20$ — reaching OR PASSING $20$ is what matters, not an exact match.
✓ the running total first passes $20$ there, at $28$ — (B) The running total reaches $4+9+15=28$ at 30-40, the first class where it passes $20$; only $13$ had accumulated through the class before.
it is the last class where the running total is below $20$ — 20-30, with a running total of only $13$, is still below $20$ — the median class is the NEXT one, 30-40, where the total first passes $20$.
it is determined by comparing class marks to the mean, not by a running total — The mean plays no role in locating the median class — only the running total of frequencies against $n/2$ does.
For the club-ages data, the median class 30-40 has $l=30$, $h=10$, $f=15$, $C=13$. The median age is
$48$ years, from using $n=40$ instead of $n/2=20$
$24.67$ years, from using the running total through 30-40 itself, $C=28$, instead of the total before it
$34.67$ years
$35$ years, the class mark of the median class, mistaken for the median
Check your answer
$48$ years, from using $n=40$ instead of $n/2=20$ — $30+((40-13)/15) \cdot 10=48$ uses the full $n=40$ instead of $n/2=20$, giving a value far outside the median class itself.
$24.67$ years, from using the running total through 30-40 itself, $C=28$, instead of the total before it — $C$ must be the total BEFORE the median class, $13$ — using $28$, the total through it, gives $30+((20-28)/15) \cdot 10=24.67$, the wrong value.
✓ $34.67$ years — (C) $30 + ((20-13)/15) \cdot 10 = 30+(7/15) \cdot 10 = 34.67$ years.
$35$ years, the class mark of the median class, mistaken for the median — $35$ is the midpoint of 30-40 — the median formula shifts slightly off that midpoint, to $34.67$ years.
A second club has 60 members with the same class intervals, and the running total first reaches $35$ (passing $n/2=30$) at the class 30-40, where $C=22$ and $f=20$. What is the median age?
$49$ years, from using $n=60$ instead of $n/2=30$
$34$ years
$24$ years, from using the running total through 30-40 itself, $C=35$, instead of the total before it
$35$ years, the class mark of the median class, mistaken for the median
Check your answer
$49$ years, from using $n=60$ instead of $n/2=30$ — $30+((60-22)/20) \cdot 10=49$ uses the full $n=60$ instead of $n/2=30$, giving a value far outside the median class itself.
✓ $34$ years — (B) $30 + ((30-22)/20) \cdot 10 = 30+(8/20) \cdot 10 = 34$ years.
$24$ years, from using the running total through 30-40 itself, $C=35$, instead of the total before it — $C$ must be the total BEFORE the median class, $22$ — using $35$, the total through it, gives $30+((30-35)/20) \cdot 10=24$, the wrong value.
$35$ years, the class mark of the median class, mistaken for the median — $35$ is the midpoint of 30-40 — the median formula shifts slightly off that midpoint, to $34$ years.
The median class is not read off the frequency column at all: it is the first class whose running total has got past halfway.
Worked example
Median age of fitness-club members
classes 10–20 to 50–60 years; frequencies $4, 9, 15, 8, 4$; $n = 40$ the frequency table for 40 fitness-club members’ ages
$n/2 = 20$; running total $4, 13, 28, 36, 40$ add frequencies class by class to build the running total
running total first reaches $20$ at $28$, in the class 30–40 — the median class the first class where the running total is at least $n/2$
$l = 30$, $h = 10$, $f = 15$, $C = 13$ $C$ is the running total through the class before the median class. $f$ is the median class’s own frequency.
median $= 30 + ((20 - 13)/15) \cdot 10 = 30 + (7/15) \cdot 10 \approx 34.67$ years substitute into the median formula for grouped data
$30 < 34.67 < 40$ Check the size of the answer. The median must sit inside the median class it came from, never outside it, and here it does.
Find the median of grouped data.
Find n over two Add all the frequencies to get $n$, then halve it. For the fitness-club table, $n = 40$, so $n/2 = 20$.
Build the running total Add the frequencies class by class. $4, 13, 28, 36, 40$.
Find the median class Find the first running total that reaches $n/2 = 20$. That is $28$, in the class 30–40.
Read l, h, f and C Read the median class’s lower limit and class size, its own frequency, and the running total before it. $l = 30$, $h = 10$, $f = 15$, $C = 13$.
Substitute Put the five numbers into the median formula. $30 + ((20 - 13)/15) \cdot 10 \approx 34.67$ years.
Check the size The median must sit inside the median class, between $30$ and $40$. $34.67$ does.
This chapter’s recap ties the mean, mode and median back to one shared object. What is it?
three separate frequency tables, one for each measure
the same list of raw, ungrouped values
the same class mark for all three measures
the same grouped frequency table
Check your answer
three separate frequency tables, one for each measure — Every worked example in this chapter reuses the SAME frequency table for whichever measure it computes — no separate table is built for each one.
the same list of raw, ungrouped values — Once data is grouped, the individual raw values are no longer what the formulas read — they read the class frequencies in the grouped table.
the same class mark for all three measures — The mean is built from class marks, but the median and mode formulas centre on the median/modal class’s lower limit and its neighbouring or running-total frequencies instead.
✓ the same grouped frequency table — (D) All three measures start from one frequency table; they simply read different information out of it.
The mean has shortcut methods (assumed-mean, step-deviation) to ease hand computation. Why do the mode and median formulas need no equivalent shortcut?
because the mode and median are always simpler ideas than the mean
because the mode and median formulas do not use the class size $h$ at all
because a shortcut for the mode and median already exists but this chapter simply omits it
they already use small numbers, from one or two classes
Check your answer
because the mode and median are always simpler ideas than the mean — Whether an idea is conceptually simple is beside the point — shortcuts exist to reduce ARITHMETIC load, and the mode/median formulas never carry the mean’s class-by-class summation.
because the mode and median formulas do not use the class size $h$ at all — Both the mode formula and the median formula multiply their correction fraction by $h$ — $h$ is very much part of both.
because a shortcut for the mode and median already exists but this chapter simply omits it — No shortcut method for the mode or median is part of this chapter’s syllabus — there is nothing omitted, because there is nothing to shorten.
✓ they already use small numbers, from one or two classes — (D) The mode and median formulas plug in a handful of numbers from one or two classes; the mean is the only one summing a weighted term over every class in the table.
A student has already found the mean and the mode of a grouped table but not the median. Which piece of information, already used for the mode, is reused when finding the median?
the class size $h$, common to both formulas
the modal class itself, which is always identical to the median class
the frequencies $f_0$ and $f_2$, which the median formula reuses unchanged
nothing is reused — the mode and median formulas share no common quantity
Check your answer
✓ the class size $h$, common to both formulas — (A) Both the mode and median formulas multiply a correction fraction by the class size $h$ — but each still needs its own class’s specific numbers besides that.
the modal class itself, which is always identical to the median class — The modal class and the median class are found by two different rules — highest frequency for one, running total against $n/2$ for the other — and can be different classes entirely.
the frequencies $f_0$ and $f_2$, which the median formula reuses unchanged — The median formula never uses $f_0$ or $f_2$ at all — it uses $C$, the cumulative frequency before the median class, and $f$, the median class’s own frequency.
nothing is reused — the mode and median formulas share no common quantity — Both formulas multiply their correction fraction by the class size $h$ — when classes are equal in size, that much genuinely carries over.
Check yourself: the whole chapter
This chapter’s methods for the mean, median and mode of grouped data all start from
each raw value on its own, listed one by one, just as in ungrouped data
the number of classes in the table, regardless of their frequencies
a rough estimate, since grouped data can never give an exact value
a frequency table of class intervals, not the raw values one by one
Check your answer
each raw value on its own, listed one by one, just as in ungrouped data — Grouped-data formulas never look at the individual values inside a class — only the class frequencies, through the class mark.
the number of classes in the table, regardless of their frequencies — The number of classes alone says nothing about the answer — the frequencies inside each class are what every formula actually uses.
a rough estimate, since grouped data can never give an exact value — These formulas give one definite, computed value from the frequency table — not a rough estimate.
✓ a frequency table of class intervals, not the raw values one by one — (D) Every formula in this chapter starts from a class’s frequency in a frequency table, not the individual values inside that class.
A frequency table has two classes with class marks $10$ and $20$, and frequencies $2$ and $3$. Using the direct method, what is the mean?
$(10+20)/2 = 15$, the plain average of the two class marks, ignoring the frequencies entirely
$(2 \cdot 10 + 3 \cdot 20)/(2+3) = 16$, each class mark weighted by its own frequency
$(2 \cdot 10 + 3 \cdot 20)/2 = 40$, the weighted sum divided by the number of classes instead of the total frequency
$2 \cdot 10 + 3 \cdot 20 = 80$, the weighted sum, without dividing by anything
Check your answer
$(10+20)/2 = 15$, the plain average of the two class marks, ignoring the frequencies entirely — $(10+20)/2=15$ treats the two classes as equally important, but the direct method weights each class mark by how many observations actually fall in it.
✓ $(2 \cdot 10 + 3 \cdot 20)/(2+3) = 16$, each class mark weighted by its own frequency — (B) The direct method weights each class mark by its own frequency — $\overline{x} = (2 \cdot 10 + 3 \cdot 20)/(2+3) = 80/5 = 16$.
$(2 \cdot 10 + 3 \cdot 20)/2 = 40$, the weighted sum divided by the number of classes instead of the total frequency — The denominator must be the TOTAL frequency, $2+3=5$, not the number of classes, $2$ — dividing by $2$ overstates the mean.
$2 \cdot 10 + 3 \cdot 20 = 80$, the weighted sum, without dividing by anything — $80$ is only the weighted sum — the direct-method formula still needs this divided by the total frequency, $5$, to give the mean.
The step-deviation method sets $u_i = (x_i-a)/h$, one step further than the assumed-mean method’s $d_i = x_i - a$. Compared to the assumed-mean method, step-deviation is most useful when
every class has the exact same frequency, the condition step-deviation is sometimes assumed to need
the data has an odd number of classes
the deviations $d_i = x_i - a$ all share the class size $h$ as a common factor
the assumed mean $a$ is chosen as the largest class mark in the table
Check your answer
every class has the exact same frequency, the condition step-deviation is sometimes assumed to need — Whether every class has the same frequency has nothing to do with step-deviation — the relevant sharing is among the deviations $d_i$, not the frequencies.
the data has an odd number of classes — The number of classes in the table plays no role in when step-deviation helps — only whether the deviations share $h$ as a factor.
✓ the deviations $d_i = x_i - a$ all share the class size $h$ as a common factor — (C) Step-deviation only pays off when every deviation $d_i$ shares the class size $h$ as a common factor — dividing it out leaves the smallest numbers of the three methods to multiply by hand.
the assumed mean $a$ is chosen as the largest class mark in the table — Step-deviation works for any valid choice of assumed mean $a$ — it is not restricted to picking the largest class mark.
A median class has $l=20$, $h=10$, $f=8$, $n=40$, and $C=12$ (the running total of frequencies before this class). What is the median?
$((n/2-C)/f) \cdot h = 10$, without adding $l$
$l + (n/2-C)/f = 21$, without multiplying by $h$
$l + ((n-C)/f) \cdot h = 55$, using $n$ instead of $n/2$
$l + ((n/2-C)/f) \cdot h = 30$
Check your answer
$((n/2-C)/f) \cdot h = 10$, without adding $l$ — $10$ is only the correction distance INTO the median class — the class’s lower limit, $l=20$, must still be added to locate the median itself.
$l + (n/2-C)/f = 21$, without multiplying by $h$ — $21$ adds $l$ but never multiplies the fraction by the class size $h=10$ — the correction distance is $((20-12)/8) \cdot 10=10$, not $1$.
$l + ((n-C)/f) \cdot h = 55$, using $n$ instead of $n/2$ — $55$ uses the full $n=40$ where the formula needs $n/2=20$ — the fraction becomes far too large.
✓ $l + ((n/2-C)/f) \cdot h = 30$ — (D) $n/2 = 20$, so the median $= 20 + ((20-12)/8) \cdot 10 = 20 + 10 = 30$.
A modal class has $l=10$, $h=5$, $f_1=10$, $f_0=6$, $f_2=5$. What is the mode?
$((f_1-f_0)/(2f_1-f_0-f_2)) \cdot h = 2.22$, without adding $l$
$l + ((f_1-f_0)/(2f_1-f_0-f_2)) \cdot h = 12.22$
$l + ((f_1-f_0)/(f_1-f_0-f_2)) \cdot h = -10$, using the wrong denominator
$l + ((f_1-f_0)/(2f_1-f_0-f_2)) \cdot h = 12.78$, with $f_0$ and $f_2$ swapped
Check your answer
$((f_1-f_0)/(2f_1-f_0-f_2)) \cdot h = 2.22$, without adding $l$ — $2.22$ is only the correction distance INTO the modal class — the lower limit, $l=10$, must still be added to locate the mode itself.
$l + ((f_1-f_0)/(f_1-f_0-f_2)) \cdot h = -10$, using the wrong denominator — The denominator must be $2f_1-f_0-f_2=20-6-5=9$, not $f_1-f_0-f_2=10-6-5=-1$ — using the wrong denominator even gives an impossible negative mode.
$l + ((f_1-f_0)/(2f_1-f_0-f_2)) \cdot h = 12.78$, with $f_0$ and $f_2$ swapped — $f_0$ must come from the row immediately BEFORE the modal class and $f_2$ from the row immediately after — swapping them, $f_0=5$, $f_2=6$, gives a different, wrong fraction.
Placing the assumed mean at the centre of a symmetric table makes the deviations cancel, and the correction term vanishes.
On this skewed income table, the mean (24.25 thousand rupees) and the median (20 thousand rupees) sit well apart.
MEAN, DIRECT: $\overline{x} = (\sum f_i x_i)/(\sum f_i)$. For the homework-time table, $440/20 = 22$ minutes.
MEAN, ASSUMED-MEAN OR STEP-DEVIATION: $\overline{x} = a + (\sum f_i d_i)/(\sum f_i)$. For the tea-packet table, $150 + 40/50 = 150.8$ grams.
MEDIAN: $l + ((n/2 - C)/f) \cdot h$. For the fitness-club table, $30 + (7/15) \cdot 10 \approx 34.67$ years.
CLASS MARK: the average of a class’s lower and upper limits. For $10$–$25$, $(10 + 25)/2 = 17.5$.
Every problem here starts from one grouped frequency table, and each formula reads off a different summary.
The mean multiplies each class mark by its frequency, directly or through the assumed-mean or step-deviation shortcut. Checking any one against another always gives agreement. The mode locates the modal class and its two neighbours. The median locates the median class from a running total, then applies one formula.
All three generalise the same three ideas we already met for ungrouped data: average, most frequent, and middle value. Grouping the data into class intervals only changes where the numbers come from, never what the three questions ask.
A batting average is the height the five scores would all have if the runs were shared out evenly, which is what the mean means.
You boil a list of numbers down to one figure more often than you think. Here are eight places a mean, a median or a mode gives you that number.
Working out a batting average. A batter scores $45$, $60$, $32$, $78$ and $85$ runs in $5$ innings. The maths: a batting average is the mean, the sum of the scores over the number of innings. The mean is $(45 + 60 + 32 + 78 + 85)/5 = 300/5 = 60$ runs.
Finding the median salary at a small company. A company’s $7$ employees earn, in thousands of rupees, $18, 20, 22, 25, 30, 45, 90$. The maths: the median is the middle salary once the list is in order. A high earner does not pull it up, unlike the mean. With $7$ values in order, the median is the $4$th value, $25$ thousand rupees.
The median depends only on position, so one extreme value cannot move it; the mean depends on every value, so it can.
Finding a class’s average test score. A class of $30$ students is grouped by marks. $2$ scored $0$–$20$, $8$ scored $20$–$40$, $12$ scored $40$–$60$, $6$ scored $60$–$80$ and $2$ scored $80$–$100$. The maths: the mean uses each class’s mark, weighted by how many students fall in it. Using class marks $10, 30, 50, 70, 90$, the mean is $(2 \cdot 10 + 8 \cdot 30 + 12 \cdot 50 + 6 \cdot 70 + 2 \cdot 90)/30 \approx 48.67$.
Finding the best-selling price band in a shoe shop. A shop sells shoes in five price bands. It sells $4$ pairs in ₹$1000$–₹$1200$ and $10$ in ₹$1200$–₹$1400$. It also sells $18$ in ₹$1400$–₹$1600$, $9$ in ₹$1600$–₹$1800$ and $5$ in ₹$1800$–₹$2000$. The maths: the best-selling price sits in the modal class, and the formula narrows it to one number inside that band. With $l = 1400$, $h = 200$, $f_1 = 18$, $f_0 = 10$, $f_2 = 9$, the mode is $1400 + (8/17) \cdot 200 \approx 1494.12$.
Finding the median water bill in a colony. A colony has $50$ homes grouped by monthly water bill. $6$ pay ₹$400$–₹$600$, $14$ pay ₹$600$–₹$800$ and $16$ pay ₹$800$–₹$1000$. $10$ pay ₹$1000$–₹$1200$ and $4$ pay ₹$1200$–₹$1400$. The maths: the median’s value comes from the median class’s limit, its running total and how many homes are in it. The running total passes $25$ at ₹$800$–₹$1000$, so $800 + ((25 - 20)/16) \cdot 200 = 862.5$.
Finding the average rent in an apartment block. Of $25$ flats, monthly rent runs in thousands of rupees. $3$ pay $15$–$20$, $7$ pay $20$–$25$ and $10$ pay $25$–$30$. $4$ pay $30$–$35$ and $1$ pays $35$–$40$. The maths: when class marks are large, an assumed mean turns them into smaller deviations before averaging. Taking $a = 27.5$ thousand rupees, the deviations give $27.5 + (-35)/25 = 26.1$ thousand rupees.
Finding the average daily fare collected by a cab driver. Over $50$ days, fares fell in five bands. ₹$200$–₹$300$ on $5$ days, ₹$300$–₹$400$ on $12$ days and ₹$400$–₹$500$ on $20$ days. ₹$500$–₹$600$ on $9$ days and ₹$600$–₹$700$ on $4$ days. The maths: dividing every deviation by the class size first keeps the numbers small to add. Taking $a = 450$ and $h = 100$, the mean is $450 + 100 \cdot (-5)/50 = 440$.
Setting a tiffin service’s default roti count. A tiffin service notes how many rotis each of $8$ customers orders: $2, 3, 2, 4, 2, 3, 2, 1$. The maths: the count ordered most often is the mode, so it makes the best default to pack. $2$ rotis appear $4$ times, more than any other count, so the mode is $2$.
Your turn. A shop’s daily sales, in rupees, over $7$ days are $2000, 1500, 1800, 2200, 1700, 2500, 1600$. What is the median daily sale? Answer: In order: ₹$1500$, ₹$1600$, ₹$1700$, ₹$1800$, ₹$2000$, ₹$2200$, ₹$2500$. The middle value is ₹$1800$.
practice The number of coconuts picked from each of $30$ palms in an orchard is grouped as 0–10, 10–20, 20–30, 30–40 and 40–50, with frequencies $2, 5, 10, 8, 5$. Find the mean number of coconuts per palm by the direct method. (Worked in full below — read it, then do the next two the same way.)
practice The number of seats filled in each of $40$ buses leaving a depot is grouped as 0–20, 20–40, 40–60, 60–80 and 80–100, with frequencies $5, 9, 12, 10, 4$. Find the mean number of seats filled by the assumed-mean method, taking $a = 50$. (The same three steps as above, with deviations from $a$ in place of the class marks.)
practice The water used in a day by each of $50$ households (in litres) is grouped as 0–50, 50–100, 100–150, 150–200 and 200–250, with frequencies $7, 9, 16, 12, 6$. Find the mean water used by the step-deviation method, taking $a = 125$ and $h = 50$. (Divide each deviation by $h$ before multiplying, and the last line puts the $h$ back.)
practice The daily wages (in ₹) of 25 workers are grouped as 100–120, 120–140, 140–160, 160–180 and 180–200, with frequencies $4, 6, 8, 5, 2$. Find the mean daily wage by the direct method.
practice The marks obtained by 30 students in a test are grouped as 10–20, 20–30, 30–40, 40–50 and 50–60, with frequencies $3, 7, 12, 5, 3$. Find the mean marks by the assumed-mean method, taking $a = 35$.
practice The electricity bills (in ₹) of 40 households are grouped as 0–100, 100–200, 200–300, 300–400 and 400–500, with frequencies $6, 10, 14, 7, 3$. Find the mean bill by the step-deviation method, taking $a = 250$ and $h = 100$.
practice The daily screen time (in minutes) of a group of students is grouped as 0–10, 10–20, 20–30, 30–40 and 40–50, with frequencies $5, k, 10, 7, 3$. If the mean screen time is $22$ minutes, find $k$ and the total number of students.
practice The marks scored by 20 students in an assignment are grouped as 0–5, 5–10, 10–15, 15–20 and 20–25, with frequencies $2, 4, 8, 4, 2$. Find the mean marks by the direct method.
practice The ages (in years) of 45 employees at a factory are grouped as 20–25, 25–30, 30–35, 35–40 and 40–45, with frequencies $6, 12, 15, 8, 4$. Find the mean age by the step-deviation method, taking $a = 32.5$ and $h = 5$.
Answers
Class marks $5, 15, 25, 35, 45$. $\sum f_i x_i = 840$ and $\sum f_i = 30$, so the mean is $840/30 = 28$ coconuts — between the smallest class mark and the largest, as it must be.
Deviations $-40, -20, 0, 20, 40$ about $a = 50$. $\sum f_i d_i = -20$ and $\sum f_i = 40$, so the mean is $50 + (-20/40) = 49.5$ seats.
Step deviations $-2, -1, 0, 1, 2$ about $a = 125$ with $h = 50$. $\sum f_i u_i = 1$ and $\sum f_i = 50$, so the mean is $125 + (1/50)(50) = 126$ litres.
Class marks $110, 130, 150, 170, 190$. $\sum f_i x_i = 3650$, $\sum f_i = 25$. Mean $= 3650/25 = 146$ rupees.
Class marks $15, 25, 35, 45, 55$; deviations $-20, -10, 0, 10, 20$ from $a = 35$. $\sum f_i d_i = -20$, $\sum f_i = 30$. Mean $= 35 + (-20/30) = 34.33$.
The mean equation gives $655 + 15k = 22(25 + k)$, so $7k = 105$ and $k = 15$. Total students $= 25 + k = 40$.
Class marks $2.5, 7.5, 12.5, 17.5, 22.5$. $\sum f_i x_i = 250$, $\sum f_i = 20$. Mean $= 250/20 = 12.5$.
Class marks $22.5, 27.5, 32.5, 37.5, 42.5$; $u_i = -2, -1, 0, 1, 2$ from $a = 32.5$. $\sum f_i u_i = -8$, $\sum f_i = 45$. Mean $= 32.5 + 5 \cdot (-8/45) \approx 31.61$.
Four bars are solid and one is drawn as a dashed outline, its height still unknown, with the mean marked between two of the classes.
Exercise 13.1 — further practice
practice The number of books read in a year by 20 students of a class are grouped as 0–4, 4–8, 8–12, 12–16 and 16–20 books, with frequencies $3, 5, 6, 4, 2$. Find the mean number of books read, by the direct method.
practice The maximum daily temperature (in degrees Celsius) recorded over 40 days is grouped as 15–20, 20–25, 25–30, 30–35 and 35–40, with frequencies $5, 8, 15, 8, 4$. Find the mean temperature by the assumed-mean method, taking $a = 27.5$.
practice The number of minutes 30 patients waited at a clinic is grouped as 0–10, 10–20, 20–30, 30–40 and 40–50, with frequencies $4, 7, 10, 6, 3$. Find the mean waiting time by the step-deviation method, taking $a = 25$ and $h = 10$.
practice The number of customers visiting a shop each day over 20 days is grouped as 10–20, 20–30, 30–40 and 40–50, with frequencies $5, 8, 4, 3$. What is the mean number of customers, by the direct method?
$25$
$27.5$
$30$
$32.5$
practice The daily rainfall (in millimetres) recorded over 30 days in a town is grouped as 0–8, 8–16, 16–24, 24–32 and 32–40, with frequencies $3, 6, 10, 7, 4$. Find the mean rainfall by the assumed-mean method, taking $a = 20$.
practice The speeds (in kilometres per hour) of 40 vehicles crossing a toll plaza are grouped as 40–50, 50–60, 60–70, 70–80 and 80–90, with frequencies $4, 9, 15, 8, 4$. Find the mean speed by the step-deviation method, taking $a = 65$ and $h = 10$.
practice The monthly savings (in ₹) of 50 employees are grouped as 1000–1500, 1500–2000, 2000–2500, 2500–3000 and 3000–3500, with frequencies $6, 12, 17, 10, 5$. Find the mean saving by the direct method, then check it by the step-deviation method, taking $a = 2250$ and $h = 500$.
practice The number of minutes spent on homework each day by a group of students is grouped as 20–30, 30–40, 40–50, 50–60 and 60–70, with frequencies $8, 12, k, 10, 6$. If the mean homework time is $43.8$ minutes, find $k$.
practice The number of minutes spent exercising daily by 75 members of a gym is grouped as 0–10, 10–20, 20–30, 30–40, 40–50 and 50–60, with frequencies $8, 12, p, 20, q, 5$. If the mean exercise time is $31$ minutes, find $p$ and $q$.
Answers
Class marks $2, 6, 10, 14, 18$. $\sum f_i x_i = 188$, $\sum f_i = 20$. Mean $= 188/20 = 9.4$ books.
Class marks $17.5, 22.5, 27.5, 32.5, 37.5$; deviations $-10, -5, 0, 5, 10$ from $a = 27.5$. $\sum f_i d_i = -10$, $\sum f_i = 40$. Mean $= 27.5 + (-10/40) = 27.25$ degrees Celsius.
practice The vegetables sold in a day at a stall (in kilograms), recorded over $50$ days, are grouped as 10–20, 20–30, 30–40, 40–50 and 50–60, with frequencies $5, 12, 20, 8, 5$. Find the modal weight sold. (Worked in full below — read it, then do the next two the same way.)
practice The number of trips made in a day by each of $40$ auto-rickshaw drivers is grouped as 5–10, 10–15, 15–20, 20–25 and 25–30, with frequencies $4, 10, 18, 6, 2$. Find the modal number of trips. (Find the modal class first, then read its two neighbours off the table.)
practice The distance walked in a day by each of $50$ postal workers (in kilometres) is grouped as 0–4, 4–8, 8–12, 12–16 and 16–20, with frequencies $8, 20, 12, 6, 4$. Find the modal distance walked. (The modal class is not the middle one here — read the frequency column before you decide.)
practice The number of goals scored by 30 football teams in a season is grouped as 0–2, 2–4, 4–6, 6–8 and 8–10, with frequencies $6, 10, 8, 4, 2$. Find the modal number of goals, and also the mean number of goals, for comparison.
practice The weekly allowance (in ₹) of 50 children is grouped as 20–40, 40–60, 60–80, 80–100 and 100–120, with frequencies $8, 14, 16, 7, 5$. Find the modal weekly allowance.
practice The number of runs scored by 35 batters in a tournament is grouped as 0–20, 20–40, 40–60, 60–80 and 80–100, with frequencies $5, 9, 12, 6, 3$. Find the modal number of runs.
practice The electricity units consumed by 40 households in a month are grouped as 0–50, 50–100, 100–150, 150–200 and 200–250 units, with frequencies $6, 9, 14, 7, 4$. Find the modal electricity consumption.
practice The number of plants in 30 gardens of a colony is grouped as 0–3, 3–6, 6–9, 9–12 and 12–15, with frequencies $4, 9, 10, 5, 2$. Find the modal number of plants, and also the mean, for comparison.
Answers
Modal class 30–40, with $l = 30$, $h = 10$, $f_1 = 20$, $f_0 = 12$ and $f_2 = 8$. Mode $= 30 + (8/20) \cdot 10 = 34$ kilograms, which lies inside 30–40.
Modal class 15–20, with $l = 15$, $h = 5$, $f_1 = 18$, $f_0 = 10$ and $f_2 = 6$. Mode $= 15 + (8/20) \cdot 5 = 17$ trips.
Modal class 4–8, the second class and not the middle one, with $l = 4$, $h = 4$, $f_1 = 20$, $f_0 = 8$ and $f_2 = 12$. Mode $= 4 + (12/20) \cdot 4 = 6.4$ kilometres.
practice The number of minutes spent on household chores daily by 50 teenagers is grouped as 0–5, 5–10, 10–15, 15–20 and 20–25, with frequencies $6, 12, 20, 8, 4$. Find the modal time spent on chores.
practice The number of late arrivals recorded over 30 days at a train station is grouped as 0–5, 5–10, 10–15, 15–20 and 20–25, with frequencies $4, 9, 11, 4, 2$. Which class is the modal class?
0–5
5–10
10–15
15–20
practice The number of books borrowed per day from a library over 18 days is grouped as 0–6, 6–12, 12–18, 18–24 and 24–30 books, with frequencies $2, 3, 9, 3, 1$. Find the modal number of books borrowed.
practice The daily temperature (in degrees Celsius) recorded over 40 days in a hill town is grouped as 20–30, 30–40, 40–50, 50–60 and 60–70, with frequencies $5, 8, 16, 8, 3$. Find the modal temperature, and also the mean temperature, for comparison.
practice The number of text messages sent daily by 60 teenagers is grouped as 0–20, 20–40, 40–60, 60–80 and 80–100, with frequencies $8, 15, 22, 10, 5$. Find the modal number of messages sent, correct to two decimal places.
practice The number of minutes spent on social media daily by 45 students is grouped as 0–8, 8–16, 16–24, 24–32 and 32–40, with frequencies $8, 10, 18, 6, 3$. Find the modal time spent, and also the mean time, correct to two decimal places, for comparison.
practice For a delivery-time distribution, the modal class is 20–30 minutes, with lower limit $l = 20$, class size $h = 10$, modal-class frequency $f_1 = 25$, and the frequency of the class after it, $f_2 = 15$. If the mode works out to $25$ minutes, find $f_0$, the frequency of the class before the modal class.
practice The heights of $50$ saplings in a nursery (in centimetres) are grouped as 10–20, 20–30, 30–40, 40–50 and 50–60, with frequencies $7, 8, 20, 10, 5$. Find the median height. (Worked in full below — read it, then do the next two the same way.)
practice The number of tomatoes picked from each of $60$ plants is grouped as 0–20, 20–40, 40–60, 60–80 and 80–100, with frequencies $10, 12, 16, 14, 8$. Find the median number of tomatoes. (Running totals first, then the class where they pass $n/2$.)
practice The masses of $40$ guavas from an orchard (in grams) are grouped as 100–120, 120–140, 140–160, 160–180 and 180–200, with frequencies $5, 7, 16, 8, 4$. Find the median mass. (Take $C$ from the class before the median class, never from the median class itself.)
practice The weights (in kg) of 40 wrestlers are grouped as 40–45, 45–50, 50–55, 55–60 and 60–65, with frequencies $4, 10, 14, 8, 4$. Find the median weight.
practice The daily sales (in ₹ thousand) of 50 shops are grouped as 0–10, 10–20, 20–30, 30–40 and 40–50, with frequencies $6, 11, 18, 10, 5$. Find the median daily sales.
practice The number of hours of sleep reported by 45 students is grouped as 4–5, 5–6, 6–7, 7–8 and 8–9 hours, with frequencies $3, 8, 20, 10, 4$. Find the median sleep time.
practice The monthly electricity bills (in ₹) of 60 households are grouped as 100–150, 150–200, 200–250, 250–300 and 300–350, with frequencies $8, 14, 20, 12, 6$. Find the median bill, and also the mean and the modal bill, comparing all three.
practice The daily temperature (in °C) recorded over 35 days is grouped as 20–24, 24–28, 28–32, 32–36 and 36–40, with frequencies $4, 8, 13, 7, 3$. Find the median temperature.
Answers
Running totals $7, 15, 35, 45, 50$ and $n/2 = 25$, so the median class is 30–40, with $l = 30$, $h = 10$, $f = 20$ and $C = 15$. Median $= 30 + (10/20) \cdot 10 = 35$ centimetres.
Running totals $10, 22, 38, 52, 60$ and $n/2 = 30$, so the median class is 40–60, with $l = 40$, $h = 20$, $f = 16$ and $C = 22$. Median $= 40 + (8/16) \cdot 20 = 50$ tomatoes.
Running totals $5, 12, 28, 36, 40$ and $n/2 = 20$, so the median class is 140–160, with $l = 140$, $h = 20$, $f = 16$ and $C = 12$. Median $= 140 + (8/16) \cdot 20 = 150$ grams.
$n = 40$, $n/2 = 20$. Running total $4, 14, 28, 36, 40$ first reaches $20$ at $28$, class 50–55. $l = 50$, $h = 5$, $f = 14$, $C = 14$. Median $= 50 + ((20 - 14)/14) \cdot 5 = 52.14$ kg.
$n = 50$, $n/2 = 25$. Running total $6, 17, 35, 45, 50$ first reaches $25$ at $35$, class 20–30. $l = 20$, $h = 10$, $f = 18$, $C = 17$. Median $= 20 + ((25 - 17)/18) \cdot 10 = 24.44$.
$n = 45$, $n/2 = 22.5$. Running total $3, 11, 31, 41, 45$ first reaches $22.5$ at $31$, class 6–7. $l = 6$, $h = 1$, $f = 20$, $C = 11$. Median $= 6 + ((22.5 - 11)/20) \cdot 1 = 6.58$ hours.
$n = 60$, $n/2 = 30$. Running total $8, 22, 42, 54, 60$ first reaches $30$ at $42$, class 200–250. $l = 200$, $h = 50$, $f = 20$, $C = 22$. Median $= 200 + ((30 - 22)/20) \cdot 50 = 220$. Mean: class marks $125, 175, 225, 275, 325$, $\sum f_i x_i = 13200$, mean $= 13200/60 = 220$. Mode: modal class 200–250, $f_1 = 20$, $f_0 = 14$, $f_2 = 12$, mode $= 200 + ((20 - 14)/(40 - 14 - 12)) \cdot 50 = 221.43$.
$n = 35$, $n/2 = 17.5$. Running total $4, 12, 25, 32, 35$ first reaches $17.5$ at $25$, class 28–32. $l = 28$, $h = 4$, $f = 13$, $C = 12$. Median $= 28 + ((17.5 - 12)/13) \cdot 4 = 29.69$ °C.
Exercise 13.3 — further practice
practice The number of minutes spent commuting daily by 60 office workers is grouped as 0–10, 10–20, 20–30, 30–40 and 40–50, with frequencies $8, 12, 20, 15, 5$. Find the median commuting time.
practice The daily screen time (in minutes) of 40 primary-school children is grouped as 0–8, 8–16, 16–24, 24–32 and 32–40, with frequencies $5, 9, 16, 7, 3$. Find the median screen time.
practice The number of visitors to a museum each day over 50 days is grouped as 0–20, 20–40, 40–60, 60–80 and 80–100, with frequencies $6, 10, 18, 11, 5$. Which class is the median class?
0–20
20–40
40–60
60–80
practice The number of pages read daily by 35 readers in a book club is grouped as 0–6, 6–12, 12–18, 18–24 and 24–30 pages, with frequencies $4, 7, 12, 9, 3$. Find the median number of pages read.
practice The waiting time (in minutes) of 50 patients at a clinic is grouped as 0–10, 10–20, 20–30, 30–40 and 40–50, with frequencies $6, 9, 20, 10, 5$. Find the median waiting time, and also the mean waiting time, for comparison.
practice The number of minutes spent stretching before practice, recorded for 50 athletes, is grouped as 0–5, 5–10, 10–15, 15–20 and 20–25, with frequencies $5, 10, 18, 12, 5$. Find the median stretching time and the modal stretching time, both correct to two decimal places, for comparison.
practice The number of minutes spent reading daily by a group of students is grouped as 0–10, 10–20, 20–30, 30–40 and 40–50, with frequencies $5, 10, 20, 12, x$. If the median reading time is $25$ minutes, find $x$.
practice The weights (in kilograms) of 50 parcels handled by a courier office are grouped as 0–10, 10–20, 20–30, 30–40, 40–50 and 50–60, with frequencies $2, 5, p, 12, q, 3$. If the median weight is $40$ kilograms, find $p$ and $q$.
Answers
$n = 60$, $n/2 = 30$. Running total $8, 20, 40, 55, 60$ first reaches $30$ at $40$, class 20–30. $l = 20$, $h = 10$, $f = 20$, $C = 20$. Median $= 20 + ((30 - 20)/20) \cdot 10 = 25$ minutes.
$n = 40$, $n/2 = 20$. Running total $5, 14, 30, 37, 40$ first reaches $20$ at $30$, class 16–24. $l = 16$, $h = 8$, $f = 16$, $C = 14$. Median $= 16 + ((20 - 14)/16) \cdot 8 = 19$ minutes.
C — 40–60.
$n = 35$, $n/2 = 17.5$. Running total $4, 11, 23, 32, 35$ first reaches $17.5$ at $23$, class 12–18. $l = 12$, $h = 6$, $f = 12$, $C = 11$. Median $= 12 + ((17.5 - 11)/12) \cdot 6 = 15.25$ pages.
$n = 50$, $n/2 = 25$. Running total $6, 15, 35, 45, 50$ first reaches $25$ at $35$, class 20–30. $l = 20$, $h = 10$, $f = 20$, $C = 15$. Median $= 20 + ((25 - 15)/20) \cdot 10 = 25$ minutes. Mean: class marks $5, 15, 25, 35, 45$, $\sum f_i x_i = 1240$, mean $= 1240/50 = 24.8$ minutes.
The median class is 20–30 with $l = 20$, $h = 10$, $f = 20$, $C = 15$ (fixed, since $x$ sits in the class after it). The median equation gives $n/2 = 25$, so $n = 50$ and $x = 50 - 47 = 3$.
The frequencies give $p + q = 28$. The median class is 30–40 with $l = 30$, $h = 10$, $f = 12$, $C = 7 + p$. The median equation gives $p = 6$, then $q = 28 - 6 = 22$.
The median formula reads a straight line between two running totals, so 25 sits where the climb from 20 to 40 crosses half of 60.
Two similar frequency tables give a mode and a median that sometimes land on the same class and sometimes land on different ones.