A comprehensive guide to the three fundamental measures of central tendency, including calculation methods, frequency tables, grouped data, missing values, comparisons, real-life applications, common errors and deeper interpretation.
When a collection of numbers is large, it can be difficult to understand it by looking at every value separately. Measures of central tendency provide useful ways of describing where the data are centred. The three most familiar measures are the mean, median and mode. They do not always give the same answer, because each describes the data in a different way. Choosing the most appropriate measure depends on the nature of the data and the question being asked.
Central tendency describes a typical or central value in a collection of data. Rather than listing every observation, we can use one representative value to summarize the data.
| Measure | Basic idea | Main calculation |
|---|---|---|
| Mean | Equal-share average | Add all values and divide by the number of values |
| Median | Middle position | Order the data and identify the centre |
| Mode | Most frequent value | Find the value occurring most often |
The three measures answer slightly different questions. If someone asks for the “average”, they may mean the mean, but it is important to determine exactly which measure is required.
The mean is found by adding all the values and dividing by how many values there are.
Another notation is:
Here, Σx means the sum of the values and n means the number of values.
Find the mean of 6, 8, 10, 12 and 14.
Sum = 6 + 8 + 10 + 12 + 14 = 50.
Number of values = 5.
Mean = 50 ÷ 5 = 10.
The mean does not have to be one of the original values. In the example above it happens to be one, but a mean such as 10.4 can be perfectly valid even when none of the observations is 10.4.
A useful way to understand the mean is to imagine that all the data are combined and then shared equally.
Suppose five containers contain 4, 6, 8, 10 and 12 units. There are 40 units altogether. If the 40 units are shared equally among five containers, each receives 8 units. Therefore the mean is 8.
This interpretation helps explain why the mean is often called an average: it represents the equal amount each observation would have if the total were redistributed evenly.
The median is the middle value when the data are arranged in order.
The most important rule is simple: sort the data first. If the values are not in order, the apparent middle number may not be the median.
When there is an odd number of values, there is one value exactly in the middle.
Data: 9, 3, 7, 5, 11.
Ordered data: 3, 5, 7, 9, 11.
There are five values, so the third value is the middle value. Median = 7.
When there is an even number of values, there are two central values. The median is the mean of those two values.
Data: 4, 12, 7, 9, 15, 5.
Ordered data: 4, 5, 7, 9, 12, 15.
The two middle values are 7 and 9.
Median = (7 + 9) ÷ 2 = 8.
For n ordered values:
For example, with 9 values the median is at position (9 + 1) ÷ 2 = 5. With 10 values, the middle positions are 5 and 6.
The mode is the value that occurs most frequently.
Data: 2, 4, 4, 5, 7, 4, 9.
The value 4 occurs three times. Every other value occurs fewer times.
Mode = 4.
Mode is particularly useful when the most common category or value is important.
Yes. If two different values share the highest frequency, the data are bimodal. If more than two values share the highest frequency, the data may be described as multimodal.
Data: 2, 2, 4, 5, 5, 7, 8.
Both 2 and 5 occur twice, while the other values occur once.
Modes = 2 and 5.
If every value occurs exactly once, there is no mode.
| Feature | Mean | Median | Mode |
|---|---|---|---|
| Uses every value? | Yes | No | No |
| Requires ordered data? | No | Yes | No |
| Can have more than one? | No, one arithmetic mean | One central result | Yes |
| Affected strongly by extreme values? | Yes | Usually much less | Usually not |
| Can be used for categories? | Only numerical data | Ordered numerical data | Yes, including categories |
Consider the data: 2, 3, 3, 4, 8.
Mean = (2 + 3 + 3 + 4 + 8) ÷ 5 = 20 ÷ 5 = 4.
Median = 3, the middle value.
Mode = 3, because it occurs twice.
The three answers are different because they measure different aspects of the same data.
The mean uses every value, so an unusually large or small observation can pull it toward itself.
Data set A: 10, 11, 12, 13, 14.
Mean = 12 and median = 12.
Now replace 14 with 100: 10, 11, 12, 13, 100.
Mean = 146 ÷ 5 = 29.2, while median remains 12.
This demonstrates why the median can be a better description of the centre when data contain extreme values.
In a perfectly symmetrical distribution with one central peak, the mean, median and mode can coincide.
Data: 2, 3, 4, 4, 4, 5, 6.
Mean = 28 ÷ 7 = 4.
Median = 4.
Mode = 4.
When the three measures are close together, the data may have a fairly balanced centre. However, their relationship alone does not prove that a distribution has a particular shape.
When a distribution has a long tail toward one side, the mean may be pulled toward that tail. The median is usually less affected.
This is one reason analysts often report the median rather than the mean for strongly skewed data.
| Situation | Often useful measure | Why |
|---|---|---|
| Numerical data without strong extremes | Mean | Uses every observation |
| Numerical data with extreme values | Median | Less affected by extremes |
| Most common size, colour or category | Mode | Identifies the most frequent item |
| Data are categorical | Mode | Mean and median may not be meaningful |
There is no single measure that is always “best”. The appropriate choice depends on the data and the purpose of the summary.
When values repeat, a frequency table provides a more efficient way to calculate the mean.
Here f is the frequency and x is the value.
| Value x | Frequency f | fx |
|---|---|---|
| 2 | 3 | 6 |
| 4 | 2 | 8 |
| 6 | 4 | 24 |
| 8 | 1 | 8 |
| Total | 10 | 46 |
Mean = Σfx ÷ Σf = 46 ÷ 10 = 4.6.
For a frequency table, the median is found by locating the central observation in the cumulative sequence.
First calculate the total frequency. Then determine the middle position(s), and use the cumulative frequencies to locate the corresponding value.
| Value | Frequency | Cumulative frequency |
|---|---|---|
| 1 | 2 | 2 |
| 2 | 3 | 5 |
| 3 | 4 | 9 |
| 4 | 1 | 10 |
Total frequency = 10. The middle positions are 5 and 6.
The 5th observation is 2 and the 6th observation is 3.
Median = (2 + 3) ÷ 2 = 2.5.
The mode is usually the value with the highest frequency.
If a table has frequencies 3, 8, 5 and 2 for values 1, 2, 3 and 4 respectively, the highest frequency is 8.
Mode = 2.
If the mean and the number of values are known, the total sum can be recovered:
The mean of five values is 18. Four values are 12, 16, 20 and 21. Find the fifth value.
Total required = 18 × 5 = 90.
Known total = 12 + 16 + 20 + 21 = 69.
Missing value = 90 − 69 = 21.
It is often faster to work with the existing total than to recalculate everything from the beginning.
Five values have mean 12. A sixth value, 18, is added.
Original total = 5 × 12 = 60.
New total = 60 + 18 = 78.
New mean = 78 ÷ 6 = 13.
Eight values have mean 15. One value, 22, is removed.
Original total = 8 × 15 = 120.
New total = 120 − 22 = 98.
Seven values remain, so new mean = 98 ÷ 7 = 14.
When combining groups, do not simply average their means unless the groups contain the same number of observations. Use totals.
Group A has 20 observations with mean 12. Group B has 30 observations with mean 18.
Group A total = 20 × 12 = 240.
Group B total = 30 × 18 = 540.
Combined mean = (240 + 540) ÷ (20 + 30) = 780 ÷ 50 = 15.6.
A common mistake is to calculate (mean A + mean B) ÷ 2. That gives the correct combined mean only when both groups have equal sizes.
One group has 10 observations with mean 20. Another has 90 observations with mean 30.
The simple average of the means is 25, but that treats both groups as equally large.
Correct combined mean = (10×20 + 90×30) ÷ 100 = 29.
A weighted mean gives different importance or weight to different values.
Suppose three components have scores 70, 80 and 90 with weights 2, 3 and 5.
Weighted total = 70×2 + 80×3 + 90×5 = 140 + 240 + 450 = 830.
Total weight = 2 + 3 + 5 = 10.
Weighted mean = 830 ÷ 10 = 83.
When numerical data are grouped into intervals, the exact individual observations are not known. A common estimate of the mean uses the class midpoint.
The midpoint of an interval is:
For the interval 10–20, the midpoint is (10 + 20) ÷ 2 = 15. If its frequency is 4, its contribution to Σfx is 4 × 15 = 60.
Because every observation within an interval is replaced by the midpoint for this calculation, the resulting mean is an estimate, not necessarily the exact mean of the original data.
For grouped data, the exact median cannot generally be known because individual values inside an interval are not given. However, the median can be estimated using the cumulative frequency and the interval containing the middle observation.
The important ideas are the median position, median class, cumulative frequency before that class, class frequency and class width.
Here L is the lower boundary of the median class, N is total frequency, CF is cumulative frequency before the median class, f is the frequency of the median class and h is the class width.
For grouped data, the modal class is the interval with the highest frequency. The exact mode may not be known because the individual observations within that interval are unknown.
An estimated grouped-data mode can be calculated using the modal class and the frequencies of the adjacent classes:
Here L is the lower boundary of the modal class, f₁ is its frequency, f₀ is the frequency before it, f₂ is the frequency after it, and h is the class width.
These measures are widely used to summarize information:
If a shop wants to know the most commonly purchased size, the mode may be more useful than the mean. If it wants the typical spending amount and a few unusually large purchases exist, the median may provide a more representative picture than the mean.
Not every data set consists of numerical measurements. Some data are categories, such as transport type, favourite fruit or colour.
For nominal categories there may be no meaningful numerical order, so mean and median are not appropriate. The mode can still identify the most common category.
Transport choices: bus, car, bus, bicycle, bus, train, car.
The most frequent category is bus, so bus is the mode.
If the same number is added to every observation, the mean and median increase by that number, and the mode also shifts by that number if it exists.
Data: 4, 6, 8, 8, 10.
Mean = 7.2, median = 8, mode = 8.
Add 5 to every value: 9, 11, 13, 13, 15.
New mean = 12.2, median = 13, mode = 13.
If every value is multiplied by the same positive number, the mean, median and mode are multiplied by that number as well.
Data: 2, 3, 3, 5, 7. Mean = 4, median = 3, mode = 3.
Multiply every value by 2: 4, 6, 6, 10, 14.
New mean = 8, median = 6, mode = 6.
The range is not a measure of central tendency. It is a simple measure of spread.
For 5, 6, 6, 7, 16:
Mean = 8.
Median = 6.
Mode = 6.
Range = 16 − 5 = 11.
Reporting a central measure together with a measure of spread often gives a more informative description of the data.
A mean should lie between the smallest and largest values in a numerical data set. If your calculated mean falls outside that range, there is an error.
If all observations lie between 20 and 50, a mean of 57 cannot be correct. Recheck the addition, the number of observations and the division.
This simple check is valuable when calculations are long or when a frequency table is involved.
| Task | Formula / rule |
|---|---|
| Mean | Σx ÷ n |
| Mean from frequency table | Σfx ÷ Σf |
| Sum from known mean | Mean × number of values |
| Median, odd n | Value at position (n + 1) ÷ 2 after ordering |
| Median, even n | Mean of values at positions n ÷ 2 and (n ÷ 2) + 1 |
| Mode | Most frequent value/category |
| Range | Maximum − minimum |
| Weighted mean | Σ(weight × value) ÷ Σweights |
| Grouped estimated mean | Σ(f × midpoint) ÷ Σf |
| Midpoint | (lower boundary + upper boundary) ÷ 2 |
The mean, median and mode are three different tools for describing the centre of data.
The mean uses every value and represents an equal-share average. It is powerful but can be strongly affected by extreme observations.
The median is the middle value after the data have been ordered. It is especially useful when the data contain unusually high or low values.
The mode identifies the most frequently occurring value or category. It can be particularly useful for categorical information and for finding the most common choice, size or measurement.
Good data handling does not simply calculate a number. It asks what that number means, whether it is an appropriate summary, how the data are distributed, and whether unusual observations affect the conclusion. Using mean, median and mode thoughtfully turns a list of numbers into meaningful information.
This article is original EDUSAMBAM educational writing. It presents standard methods for calculating, comparing and interpreting mean, median and mode, including frequency tables, grouped-data estimates, weighted means and practical data-handling applications.
20 questions covering mean, median, mode, frequency tables, missing values, grouped data and interpretation. Answer every question, then submit to see your score instantly.