Topic (iv) — Tabulation and Compilation of Data
<- Previous · Index of all topics · Next ->
Part of the full combined notes. · 3 PYQ callouts inside.
🔷 TOPIC (iv) — TABULATION AND COMPILATION OF DATA
Syllabus topic (iv) of 10 · runs until Topic (v)
Covers: Classification of data & its principles · statistical series · time / geographical / magnitude / condition classification · arrangement of data (individual, discrete, continuous) · tabulation & parts of a table · types of tables · graphical representation — histogram, bar diagram, frequency polygon, frequency curve, ogives · true, non-true & open-end classes 🎯 Asked in the papers: 2024 · Q74 — which statement about tabulation is FALSE · 2022 · Q38 — shape of a frequency polygon as classes increase
1. Classification of Data
Process of arranging data into different groups or classes based on common characteristics.
Characteristics of classification of Data
- It should be unambiguous
- Flexible to adjustments
- It should perform homogeneous grouping
- Its basis should be clearly defined and adhered to.
Objectives
- To simplify the huge data
- To facilitate comparison.
- To provide basis for tabulation
- To make data understandable.
- Clearly specify similarities and dissimilarities.
Representation of Data
graph LR
A[Representation of Data] --- B[Text]
A --- C[Tabular]
A --- D[Graphical]
2. Principles of Classification
**1. Exhaustive **
- Every item should be classified.
- NO Residual or miscellaneous class.
**2. Mutually Exclusive ** Every single item should fall in only one class.
**3. Stability ** Common principle should be maintained / followed for whole classification.
**4. Suitability ** Classification should be suitable as per objective of research.
**5. Flexibility ** Classification should be adjustable to new situation / new adjustments.
**6. Homogeneity ** Data items of a class should be similar in characteristics.
3. Statistical Series
Arrangement of data in a logical order.
graph TD
Root["==Statistical Series=="]
Root --> Char["On the basis of<br>characteristics"]
Root --> Const["On the basis of<br>construction"]
Char --> C1["Time-Series"]
Char --> C2["Spatial<br>series"]
Char --> C3["Condition<br>series"]
Char --> C4["Magnitude<br>series"]
C2 -.-> OR["(Or)"]
OR -.-> Simple["Simple<br>(one attribute)"]
OR -.-> Manifold["Manifold<br>(more than<br>one attributes)"]
Const --> S1["Individual<br>series"]
Const --> S2["Discrete<br>series"]
Const --> S3["Continuous<br>series"]
4. Important Classifications
4.1 Time-Series / Chronological / Temporal
"Demand: Sep 100, Oct 200, Nov 300, Dec 400. What is the 3-month moving average for January?" a) 300 ✅ · b) 350 · c) 400 · d) 450 ⭐ Answer: A. A 3-month moving average for January = mean of the previous three months = (200 + 300 + 400) ÷ 3 = 300. ⚠️ Moving averages are NOT in the 10-topic FAA statistics syllabus — this came from the 2022 Panchayat paper. Low priority, one line to learn.
Based on time of occurence. e.g population of India from 2010 - 2015
| Year | Population |
|---|---|
| 2010 | 106 cr. |
| 2011 | 110 cr. |
| 2012 | 114 cr. |
| 2013 | 120 cr. |
| 2014 | 126 cr. |
| 2015 | 130 cr. |
4.2 Geographical / Spatial
Classification on the basis of Area/Region/Location. e.g. sales report of company
| State | Sales (Lakhs) |
|---|---|
| Delhi | 30 |
| J&K | 20 |
| Punjab | 40 |
| Kolkata | 15 |
| Hyderabad | 30 |
4.3 Magnitude / Numerical / Quantitative
Based on quantity
| Height (ft) | No. of Persons |
|---|---|
| 5.10 | 25 |
| 5.11 | 15 |
| 6 | 10 |
| 6.1 | 5 |
4.4 Condition / Seasonal
When classification is done on the basis of seasonal variations/situations. e.g sales of ice cream / cold drinks in summers and winters.
**Simple classification ** Classification on the basis of only one attribute. e.g. On the basis of Gender
graph TD
A[Gender] --> B[Male]
A --> C[Female]
**Manifold classification ** Classification based on more than one attributes.
graph TD
Pop[Population] --> Lit[Literate]
Pop --> Illit[Illiterate]
Lit --> Emp1[Employed]
Lit --> Unemp1[Unemployed]
Illit --> Emp2[Employed]
Illit --> Unemp2[Unemployed]
Emp1 --> M1[Married]
Emp1 --> UM1[Unmarried]
Unemp1 --> M2[M]
Unemp1 --> UM2[UM]
Emp2 --> M3[M]
Emp2 --> UM3[UM]
Unemp2 --> M4[M]
Unemp2 --> UM4[UM]
5. Arrangement of Data
Data can be arranged in two ways:
- Serial (or) Alphabetic order
- Ascending (or) Descending order.
e.g. Primary, secondary, Data, Questionnaire (ungrouped data) / Raw data Data, Primary, Questionnaire, Secondary.
e.g. 13, 19, 12, 10, 25, 32, 29, 37 (ungrouped / Raw data)
Arrayed Data
10, 12, 13, 19, 25, 29, 32, 37Ascending / Increasing order37, 32, 29, 25, 19, 13, 12, 10Descending / Decreasing order
5.1 Individual Series
Each item is given separate value. e.g
| Name of student | Weight in (kgs) |
|---|---|
| X | 55 |
| Y | 70 |
| Z | 60 |
| K | 55 |
5.2 Discrete Series (Discrete Frequency Distribution)
Each individual value is presented with frequency.
e.g.
| Salary of Employees | No. of Employees | (Frequency) Tally mark |
|---|---|---|
| 20,000 | 5 | $\cancel{ |
| 40,000 | 7 | $\cancel{ |
| 50,000 | 4 | $ |
| 70,000 | 2 | $ |
| Marks | No. of students | Tally mark |
|---|---|---|
| 1 | 3 | $ |
| 3 | 5 | $\cancel{ |
| 5 | 9 | $\cancel{ |
| 7 | 10 | $\cancel{ |
| 9 | 12 | $\cancel{ |
5.3 Continuous Series (Continuous Frequency Distribution)
Shows range of values of different items.
e.g.
| Marks (class interval) | No. of students |
|---|---|
| 0-5 | 5 |
| 5-10 | 10 |
| 10-15 | 12 |
| 15-20 | 7 |
0-5class interval = (Range)- lower limit
- upper limit
- class mark = mid value =
6. Tabulation of Data
"Which statement regarding tabulation is FALSE?" a) Tabulation allows presentation of complicated data · b) Tabulation is a prerequisite for diagrammatic representation · c) Statistical analysis necessarily involves tabulation · d) Tabulation facilitates comparison between rows and NOT columns ✅ ⭐ Answer: D — that is the FALSE one. A table compares across BOTH rows AND columns. The definition immediately below ("rows and columns") is exactly what kills it. ⚠️ Underline the word FALSE before you answer.
Systematic and Logical representation of numeric data in rows ( horizontal) and columns ( vertical).
Objectives (1) To simplify the complex data. (2) To bringout important features. (3) To facilitate comparison. (4) To facilitate statistical analysis. (5) To save space and time.
Limitations (1) Lack of description (2) Incapable of presenting individual items. (3) Needs ample knowledge & understanding.
General Format
Table No: <Title>
<Head note> (if any)
| Stub (Row heads) | Caption (column Headings) | Total (Rows) | |||
|---|---|---|---|---|---|
| Sub-Heads1 | Sub-Heads2 | Subheads1 | Sub-Heads2 | ||
| <-- | BODY ↓ | --> | |||
| Total (cols) | |||||
- S. Note: -
- Foot note / Note
6.1 Main Parts of a Table
(1) Table NO
- First item mentioned on top of table
- Identification and References.
(2) Title
- Second item, just above the table / by right side of T.NO.
(3) Head-note / Prefatory
- 3rd item, just above the table.
- Information about unit of data. e.g.: currency ₹ (or) $ Quintals or tonnes
- Generally given in brackets.
(4) Caption / Col-Headings / characteristic
- Top of each column.
- Explains data in column.
(5) Stub / Row Heading
- Title of horizontal rows.
(6) Body
- Numeric data.
- In cols and rows.
(7) Foot Note
- To explain non-explanatory things.
- Exceptions if any.
- Circumstances affecting data.
(8) Source Note
- Statement indicating source of data.
Table NO. 2.1 Records of Graduation students faculty wise H.Note: (Includes all registered students)
| Faculty | Ist Year col. H | 2nd Year col. H | Total (Row) | ||
|---|---|---|---|---|---|
| Boys S.H | Girls S.H | Boys S.H | Girls S.H | ||
| B.Sc | 70 | 30 | 50 | 35 | 185 |
| B.com | 50 | 20 | 30 | 30 | 130 |
| B.A | 100 | 60 | 60 | 20 | 240 |
| Total (col) | 220 | 110 | 140 | 85 | ==555== |
Source note: Admission dept. ABC college. Footnote: Late registrations are not included.
6.2 Characteristics of a Good Table
- Title compatible with objective.
- Comparable.
- Ideal size.
- Col-total, Row total, total included.
- Stubs.
- Headings.
- Simple, Economical & Attractive.
- No Abbreviations.
- No Over crowding.
6.3 Types of Tables
graph TD
Root[Types of Tables]
Root --> Obj[objectivity / purpose]
Root --> Nat[Nature / originality]
Root --> Char[characteristic / constructio]
Obj --> O1["==General Purpose=="]
Obj --> O2["==Special Purpose=="]
Nat --> N1["==Primary / original=="]
Nat --> N2["==Secondary / Derived=="]
Char --> C1["==Simple (1-way)=="]
Char --> C2["==Complex=="]
C2 --> C2_1["==Double / Two-way=="]
C2 --> C2_2["==Treble / 3-way=="]
C2 --> C2_3["==Manifold=="]
6.4 Classification of Tables
-
On the basis of "Purpose" ① General Purpose / Master Tables
- General use.
- Not meant for special purpose. e.g.: Census
② Special purpose table / Summary Table.
- Derived from general table.
- Serves special purpose.
- Useful for calculation of analytical statistics like ratio, percentage etc. e.g.: Calculating Ratio of Male:Female
-
On the basis of "Originality" (1) Original Table / Primary Table
- Data is presented in the form in which it is collected.
(2) Derived Table / Secondary Table
- Converted into any form as per requirement. e.g.: Limited columns as required.
-
As per construction / characteristics (1) Simple Table
- "One-way" table.
- Data presented based on the "one" characteristic only.
Table 1.1 Faculty-wise Number of students
| Faculties: (Attribute) | No. of Students |
|---|---|
| Science | 30 |
| Commerce | 40 |
| Arts | 60 |
| Total | 130 |
Source note. One-way Table Foot note.
(2) Complex tables
-
More than one attribute presented simultaneously.
(i) Double / Two-way Table
- Data is tabulated on the basis of two inter-related characteristics.
Table 1.2 Faculty - wise number of Male and Female students
| FACULTY | No. of Students (gender) | Total | |
|---|---|---|---|
| BOYS | GIRLS | ||
| Arts | 35 | 20 | 55 |
| Commerce | 40 | 25 | 65 |
| Science | 30 | 35 | 65 |
| Total | 105 | 80 | 185 |
S. note. F. note.
- Three-way / Trebles
- Data is tabulated on the basis of 3-inter-related characteristics.
Table 1.3 Faculty wise, (Gender & semester) list of students
| FACULTY | No. of students | Total | |||||
|---|---|---|---|---|---|---|---|
| Girls | Boys | ||||||
| Sem I | Sem II | Total(1) | Sem I | Sem II | Total(2) | ||
| Science | 15 | 20 | 35 | 20 | 50 | 70 | 105 |
| Arts | 35 | 30 | 65 | 45 | 85 | 130 | 195 |
| Commerce | 25 | 35 | 60 | 35 | 90 | 125 | 185 |
| Total | 75 | 85 | 160 | 100 | 225 | 325 | 485 |
S. note. 3-way Table F. note.
- Manifold (Higher order Table)
- Data is tabulated on the basis of large no. of interrelated characteristics.
Table 1.4 Faculty wise (UG / PG) Gender based students in each sem.
| FACULTY | Number of students | Total | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| UG | PG | ||||||||||||
| Boys | Girls | Boys | Girls | ||||||||||
| Sem I | II | Total | I | II | Total | I | II | Total | I | II | Total | ||
| Total | |||||||||||||
S. note F. note
7. Graphical Representation of Data
Presentation of statistical data on graph paper in the form of lines or curves.
Merits (i) Simplifies complex data. (ii) Helps in forecasting. (iii) Variations in the values of variables.
Demerits (i) Precise values are not shown. (ii) May give wrong predictions. (iii) Difficult to interpret.
| Graph | Vs | Diagram |
|---|---|---|
| 1. Graph paper is required. | 1. Can be drawn on plain paper. | |
| 2. Represents mathematical relationship. | 2. Does not represent any relationship. | |
| 3. Median, mode etc can be determined. | 3. Impossible through diagram. |
7.1 Procedure to Construct a Graph
y-axis (ordinate)
^
|
O₂ | O₁
|
<---------------------+---------------------> x-axis
| (abscissa)
O₃ | O₄
|
v
(i) Scale Appropriate scale. (ii) Heading Suitable and precise heading. (iii) Proper Indications More than one curve, properly differentiated. (iv) False Base line y-axis. (v) Index lines & colors. (vi) Data Table full knowledge of data. (vii) Drawing line or curve Joining the points. (viii) Paper size Appropriate size of graph paper. (ix) From L to R & B to Top Left to right & Bottom to top.
7.2 Frequency Diagram (Histogram)
Frequency distribution is represented using graphs.
Histogram Two dimensional diagram, set of rectangles with base as the intervals between class boundaries. ** Histogram is not drawn for discrete data.
Equal classes
| Marks | No. of Students |
|---|---|
| 0 - 5 | 20 |
| 5 - 10 | 30 |
| 10 - 15 | 10 |
| 15 - 20 | 40 |
| 20 - 25 | 50 |
xychart-beta
title "Histogram (Student marks)"
x-axis ["0-5", "5-10", "10-15", "15-20", "20-25"]
y-axis "Students" 0 --> 50
bar [20, 30, 10, 40, 50]
Unequal classes
| Marks | No of students |
|---|---|
| 0 - 5 | 20 |
| 5 - 15 | 30 |
| 15 - 20 | 10 |
| 20 - 30 | 40 |
| 30 - 35 | 10 |
xychart-beta
title "Histogram (Student marks - Unequal classes)"
x-axis ["0-5", "5-15", "15-20", "20-30", "30-35"]
y-axis "Students" 0 --> 40
bar [20, 30, 10, 40, 10]
7.3 Bar Diagram
- Bars with arbitrary width are drawn to represent the data.
- Drawn for both discrete and continuous data.
- Some space has to be left between consecutive bars.
e.g.: Production of Rice in J&K from 1997 - 2004
| Year | Production (tonnes) |
|---|---|
| 1997 | 10 |
| 1998 | 20 |
| 1999 | 30 |
| 2000 | 15 |
| 2001 | 20 |
| 2002 | 25 |
| 2003 | 30 |
| 2004 | 35 |
xychart-beta
title "Bar diagram: Production of Rice"
x-axis ["1997", "1998", "1999", "2000", "2001", "2002", "2003", "2004"]
y-axis "Production (tonnes)" 0 --> 40
bar [10, 20, 30, 15, 20, 25, 30, 35]
7.4 Frequency Polygon
Frequency polygon / Frequency polygon curve is constructed by joining the mid points of the frequencies presented using rectangles (histograms). ** constructed for both discrete as well as continuous series.
e.g.: Construct frequency polygon for continuous series
| Age (Yrs) | No. of students |
|---|---|
| 5-10 | 5 |
| 10-15 | 10 |
| 15-20 | 15 |
| 20-25 | 20 |
| 25-30 | 5 |
xychart-beta
title "Frequency Polygon for Age VS No. of students"
x-axis ["5-10", "10-15", "15-20", "20-25", "25-30"]
y-axis "No. of students" 0 --> 25
bar [5, 10, 15, 20, 5]
line [5, 10, 15, 20, 5]
e.g. construct Frequency polygon for discrete series
| Age. | No. of students |
|---|---|
| 4 | 10 |
| 5 | 15 |
| 6 | 20 |
| 7 | 10 |
| 8 | 5 |
| 9 | 4 |
xychart-beta
title "Frequency polygon Age vs No. of students (discrete series)"
x-axis ["1", "2", "3", "4", "5", "6", "7", "8", "9", "10"]
y-axis "No. of students" 0 --> 25
line [0, 0, 0, 10, 15, 20, 10, 5, 4, 0]
7.5 Frequency Curve
Obtained by drawing smooth freehand curve passing through the closely associated points. ** Can be drawn for both discrete as well as continuous series.
➕ Added — How a Frequency Polygon BECOMES a Frequency Curve.
Histogram → Frequency Polygon → Frequency Curve is one continuous progression:
- Histogram — bars over each class
- Frequency polygon — join the mid-points of the tops of those bars
- ⭐ Frequency curve — now make the class intervals SMALLER, so the number of classes INCREASES. Each corner of the polygon gets shallower, and the line becomes increasingly SMOOTH. In the limit (infinitely many, infinitely narrow classes) the polygon becomes the frequency curve.
[!pyq] 🎯 2022 · Q38 — "What happens to the shape of a frequency polygon as the number of classes increases?" → ⭐ It becomes increasingly smooth
Also worth knowing: the area under a frequency polygon equals the area of its histogram · a polygon can compare two or more distributions on one graph (a histogram cannot).
7.6 Types of Curves by Shape
- Normal curve: symmetric curve / Bell curve / Gaussian distribution
- U-shaped curve
- Positive skewed curve
- Negative skewed
- Bi-modal curve
- J-shaped curve
- Reverse J-shaped curve
- Mixed Multi-modal curve
7.7 Cumulative Frequency Curve (Ogive)
| Salary (in Ks) | No. of Employees (f) |
|---|---|
| 10-20 | 7 |
| 20-30 | 10 |
| 30-40 | 15 |
| 40-50 | 18 |
| 50-60 | 20 |
| 60-70 | 25 |
Less than (LTB) cumulative frequency Less than ogive curve:
| Salary | less than cumulative |
|---|---|
| Less Than 20 | 7 |
| 30 | 17 |
| 40 | 32 |
| 50 | 50 |
| 60 | 70 |
| 70 | 95 |
Greater than (GTB) cumulative frequency Greater than ogive curve:
| Salary | More than C.F. |
|---|---|
| 10 | 95 |
| 20 | 88 |
| 30 | 78 |
| 40 | 63 |
| 50 | 45 |
| More than 60 | 25 |
** intersection of Less than ogive & greater than ogive = Median
8. Class Intervals
8.1 True Class Intervals
There is no gap between the successive classes. Upper limit of each class is equal to lower limit of the succeeding class. e.g.
| Age | (No. of persons) Frequency |
|---|---|
| 10 - 20 | 7 |
| 20 - (30) | 8 |
| (30) - 40 | 9 |
| 40 - 50 | 10 |
- Class Boundaries
8.2 Non-True Classes
Upper limit of each class is not equal to lower limit of successive class. e.g
| Weight (kgs) | No. of person |
|---|---|
| 30 - (39) | 5 |
| (40) - 49 | 10 |
| 50 - 59 | 20 |
| 60 - 69 | 25 |
- class limits
8.3 Open-End and Closed-End Classes
When a class limit is missing either at lower end (first class) or at upper end (last class) or limits are not specified at both ends — open end classes. When limits are specified at both ends — closed end classes.
Examples of open-end classes
| Marks | No. of students |
|---|---|
| Below 20 | 10 |
| 20 - 30 | 12 |
| 30 - 40 | 13 |
| 40 - 50 | 15 |
| Marks | No. of students |
|---|---|
| 0 - 20 | 5 |
| 20 - 40 | 10 |
| 40 - 60 | 15 |
| Above 60 | 20 |
| Marks | No. of students |
|---|---|
| Below 20 | 7 |
| 20 - 30 | 8 |
| 30 - 40 | 9 |
| 40 - 60 | 20 |
| Above 60 | 12 |
Example of closed-end classes
| Marks | No. of students |
|---|---|
| 0 - 10 | 5 |
| 10 - 20 | 7 |
| 20 - 30 | 9 |
| 30 - 40 | 12 |
| 40 - 50 | 15 |
9. Important Points on Compilation and Tabulation
- Repository: Tables to show data in orderly manner.
- Marginal Frequency:
- Relative Frequency =
- Univariate Frequency Distribution: One variable
- Bivariate Frequency Distribution: Two variables