Every time you collect a large batch of numbers, marks scored by 200 students, daily temperatures for a year, or household expenditure for a whole city, you run into the same problem. Listing every single value separately gives you a table so long that no pattern is visible at all. This is exactly why statisticians group raw scores into class intervals before they try to make sense of them. A grouped frequency distribution takes that unwieldy pile of numbers and sorts it into a handful of ranges, so you can see at a glance how the data is spread out.
Table of Contents
- Why large datasets need grouping
- Deciding the range and number of classes
- The exclusive method of classification
- Using tally bars to count frequencies
- The inclusive method of classification
- True or actual class limits
- Why this conversion is not optional for some calculations
- Choosing the right method for your data
Why large datasets need grouping
When you have only ten or fifteen observations, an ungrouped frequency table works fine. You simply list each value and count how many times it appears. But once the dataset runs into hundreds of values, this approach breaks down. A table with 150 rows tells you almost nothing useful. Grouping the same data into, say, ten class intervals compresses that information into a shape you can actually read and interpret.
The trade-off is that grouping sacrifices some precision. Once scores are placed inside a class such as 20-30, you know how many observations fall in that range, but not their exact individual values. To handle this loss, statisticians treat the midpoint of each class as a stand-in for every score within it, since the frequency of a class is assumed to be concentrated around its centre. This convention becomes important later when you calculate the mean, median, or mode from grouped data, and it is used consistently across standard statistics teaching material for exactly this reason.
Deciding the range and number of classes
Two questions come up before you draw a single line on your table: how wide should each class be, and how many classes should there be? The starting point is always the range, calculated by subtracting the lowest value in the dataset from the highest. If the marks in a class test run from 12 to 96, the range is 84.
There is no fixed rule for how many classes to create, though the number typically falls somewhere between 5 and 30, depending on how large and varied the dataset is. Too few classes and you lose meaningful detail; too many and you are back to the same clutter you were trying to avoid. Some textbooks recommend Sturges’ rule as a rough guide, where the number of classes is close to 1 plus the base-2 logarithm of the sample size, though even this is treated as a starting estimate rather than a strict formula, since no single rule works perfectly for every dataset.
Once you have decided on the number of classes, the width of each class interval is found by dividing the range by that number. It helps to choose a width that is conveniently divisible, such as 5, 10, or 20, rather than an awkward number like 7 or 13. This keeps the class boundaries clean and easy to read on a table or graph. Every class interval in the distribution should also be of uniform width, so that frequencies across classes remain directly comparable.
The exclusive method of classification
Once the class width is fixed, you need a rule for exactly where each score belongs, especially for values that sit right on a boundary. The exclusive method solves this by making the upper limit of one class the same as the lower limit of the next. So if your classes are 10-20, 20-30, and 30-40, a score of exactly 20 is always placed in the 20-30 class, not the 10-20 class. In other words, the class 20-30 is understood to run from 20 up to, but not including, 30, which keeps the classes continuous and prevents any value from qualifying for two classes at once.
This convention is why the method is called exclusive: each class technically excludes its upper limit. It is the standard approach used when the underlying variable is continuous, such as height, weight, marks out of 100, or time taken to finish a task, because these measurements can take any value in between whole numbers. The same left-inclusive, right-exclusive convention is followed in most formal treatments of frequency distributions.
Using tally bars to count frequencies
Once your classes are set, you go through the raw data one value at a time and mark a tally against the class it belongs to. Tally marks are usually grouped in bundles of five, with the fifth mark drawn as a diagonal line across the previous four, which makes the tallies quick to count once you are done. After every observation has been tallied, you convert the tally marks in each class into a plain frequency number.
This step comes with a built-in accuracy check. Add up the frequencies of every class in your table, and the total must exactly equal the number of observations you started with. If it does not, it usually means a value was counted twice or missed altogether while tallying, and you need to recheck the raw data against the table before moving further. This simple cross-check is one of the most useful habits to build when constructing any frequency distribution by hand.
The inclusive method of classification
The inclusive method takes a different approach to class boundaries. Here, both the lower limit and the upper limit of a class belong to that same class. Classes might look like 10-19, 20-29, and 30-39, where a value of 19 stays in the first class and a value of 20 moves into the second. Notice the small gap between 19 and 20; unlike the exclusive method, the upper limit of one class does not double up as the lower limit of the next.
This method is generally preferred when you are dealing with discrete, whole-number data, values that genuinely cannot take fractions, like the number of children in a household, the number of correct answers on a quiz, or the number of workers in a factory. Since a person cannot have 19.5 children, there is no ambiguity about where a borderline score should sit, and the gap between classes causes no practical problem. This distinction between where the inclusive method is a natural fit and where it becomes awkward is well documented in comparisons of inclusive and exclusive class intervals.
The trouble starts when the inclusive method is applied to continuous measurements. If test scores can include decimals, a value like 19.5 has no obvious home in a table built from 10-19 and 20-29. This is precisely the gap that the true class method is designed to close.
True or actual class limits
The true class method, sometimes called the actual class method, treats every recorded observation as representing not a single point but a small unit of length around that point. A score written down as 20 is assumed to actually extend from 19.5 to 20.5, half a unit below and above its face value. Applying this logic to every class boundary converts an inclusive-style table into a fully continuous one, with no gaps and no overlaps between adjacent classes.
To convert an inclusive distribution into true class limits, you find half the gap between the upper limit of one class and the lower limit of the next, then subtract that amount from every lower limit and add it to every upper limit. For classes like 10-19 and 20-29, the gap between 19 and 20 is 1, so the adjustment factor is 0.5. The class 10-19 becomes 9.5-19.5, and 20-29 becomes 19.5-29.5, giving you continuous, touching boundaries throughout the table, in line with the standard conversion procedure shown in worked histogram examples from continuous test-score data.
Why this conversion is not optional for some calculations
True class limits are not just a technical formality. Several statistical procedures assume that classes are perfectly continuous, with the upper boundary of one class touching the lower boundary of the next. Drawing a proper histogram, for instance, requires bars that sit flush against each other with no visible gaps, which is only possible once you are working with true class boundaries rather than raw inclusive limits. The same continuity assumption underlies formulas for the median and certain other measures calculated from grouped data, where class boundaries feed directly into the arithmetic.
Skipping this conversion when it is needed does not just look untidy, it can quietly distort your calculations. Whether you are organising CBSE board exam scores, income data for a household survey, or lab measurements for a college project, the components used to build a proper grouped frequency distribution table, correct class intervals, matching frequencies, and continuity where it is required, all work together to keep the final analysis accurate.
Choosing the right method for your data
In practice, the choice between exclusive, inclusive, and true class methods comes down to one question: is your variable discrete or continuous? Marks that can include decimals, heights, weights, and durations call for the exclusive method or true class limits, since these values genuinely flow into each other without natural breaks. Counts of whole, indivisible units, people, objects, or events, sit comfortably in an inclusive table, since there is no in-between value to worry about.
Getting this choice right at the start saves you from having to redo an entire table later, especially if you later need to calculate the median, draw a histogram, or run any statistical test that assumes continuous classes. A little care while setting up your classes pays off every time you use that table afterwards.
What do you think? If you were grouping the marks of 60 students where scores range from 8 to 97, would you lean towards the exclusive method or convert to true class limits, and why? And can you think of a dataset from your own coursework where the inclusive method would clearly be the better fit?
References
- https://sathee.iitk.ac.in/ncert-books/ncert-books-theory/class-10/nbt-math-10/math-10-chapter-13-statistics/
- https://socialsci.libretexts.org/Courses/Saint_Mary's_College_(Notre_Dame_IN)/Social_Science_Statistics/04:_Graphing_Distributions/4.04:_Histograms
- https://www.geeksforgeeks.org/types-of-frequency-distribution/
- https://www.cuemath.com/data/class-interval/
- https://www.vedantu.com/maths/frequency-distribution-grouped
Leave a Reply