Upcoming-ExamsFinance-Account-AssistantFAA-STATISTICSNotesTopic (iv) — Tabulation and Compilation of Data

<- Previous · Index of all topics · Next ->

Statistics — syllabus topic (iv) of 10 · 🎯 PYQs from this topic: 2024 Q74 - 2022 Q38 - 2022 Q33

Part of the full combined notes. · 3 PYQ callouts inside.


🔷 TOPIC (iv) — TABULATION AND COMPILATION OF DATA

Syllabus topic (iv) of 10 · runs until Topic (v)

Covers: Classification of data & its principles · statistical series · time / geographical / magnitude / condition classification · arrangement of data (individual, discrete, continuous) · tabulation & parts of a table · types of tables · graphical representation — histogram, bar diagram, frequency polygon, frequency curve, ogives · true, non-true & open-end classes 🎯 Asked in the papers: 2024 · Q74which statement about tabulation is FALSE · 2022 · Q38shape of a frequency polygon as classes increase


1. Classification of Data

Process of arranging data into different groups or classes based on common characteristics.

Characteristics of classification of Data

  • It should be unambiguous
  • Flexible to adjustments
  • It should perform homogeneous grouping
  • Its basis should be clearly defined and adhered to.

Objectives

  • To simplify the huge data
  • To facilitate comparison.
  • To provide basis for tabulation
  • To make data understandable.
  • Clearly specify similarities and dissimilarities.

Representation of Data

graph LR
    A[Representation of Data] --- B[Text]
    A --- C[Tabular]
    A --- D[Graphical]

2. Principles of Classification

**1. Exhaustive **

  • Every item should be classified.
  • NO Residual or miscellaneous class.

**2. Mutually Exclusive ** Every single item should fall in only one class.

**3. Stability ** Common principle should be maintained / followed for whole classification.

**4. Suitability ** Classification should be suitable as per objective of research.

**5. Flexibility ** Classification should be adjustable to new situation / new adjustments.

**6. Homogeneity ** Data items of a class should be similar in characteristics.

3. Statistical Series

Arrangement of data in a logical order.

graph TD
    Root["==Statistical Series=="]
    
    Root --> Char["On the basis of<br>characteristics"]
    Root --> Const["On the basis of<br>construction"]
    
    Char --> C1["Time-Series"]
    Char --> C2["Spatial<br>series"]
    Char --> C3["Condition<br>series"]
    Char --> C4["Magnitude<br>series"]
    
    C2 -.-> OR["(Or)"]
    OR -.-> Simple["Simple<br>(one attribute)"]
    OR -.-> Manifold["Manifold<br>(more than<br>one attributes)"]
    
    Const --> S1["Individual<br>series"]
    Const --> S2["Discrete<br>series"]
    Const --> S3["Continuous<br>series"]

4. Important Classifications

4.1 Time-Series / Chronological / Temporal

🎯 CAME IN THE EXAM — 2022 · Q33 (off the FAA syllabus, but from a time-series idea)

"Demand: Sep 100, Oct 200, Nov 300, Dec 400. What is the 3-month moving average for January?" a) 300 ✅ · b) 350 · c) 400 · d) 450 ⭐ Answer: A. A 3-month moving average for January = mean of the previous three months = (200 + 300 + 400) ÷ 3 = 300. ⚠️ Moving averages are NOT in the 10-topic FAA statistics syllabus — this came from the 2022 Panchayat paper. Low priority, one line to learn.

Based on time of occurence. e.g population of India from 2010 - 2015

YearPopulation
2010106 cr.
2011110 cr.
2012114 cr.
2013120 cr.
2014126 cr.
2015130 cr.

4.2 Geographical / Spatial

Classification on the basis of Area/Region/Location. e.g. sales report of company

StateSales (Lakhs)
Delhi30
J&K20
Punjab40
Kolkata15
Hyderabad30

4.3 Magnitude / Numerical / Quantitative

Based on quantity

Height (ft)No. of Persons
5.1025
5.1115
610
6.15

4.4 Condition / Seasonal

When classification is done on the basis of seasonal variations/situations. e.g sales of ice cream / cold drinks in summers and winters.

**Simple classification ** Classification on the basis of only one attribute. e.g. On the basis of Gender

graph TD
    A[Gender] --> B[Male]
    A --> C[Female]

**Manifold classification ** Classification based on more than one attributes.

graph TD
    Pop[Population] --> Lit[Literate]
    Pop --> Illit[Illiterate]
    
    Lit --> Emp1[Employed]
    Lit --> Unemp1[Unemployed]
    
    Illit --> Emp2[Employed]
    Illit --> Unemp2[Unemployed]
    
    Emp1 --> M1[Married]
    Emp1 --> UM1[Unmarried]
    
    Unemp1 --> M2[M]
    Unemp1 --> UM2[UM]
    
    Emp2 --> M3[M]
    Emp2 --> UM3[UM]
    
    Unemp2 --> M4[M]
    Unemp2 --> UM4[UM]

5. Arrangement of Data

Data can be arranged in two ways:

  1. Serial (or) Alphabetic order
  2. Ascending (or) Descending order.

e.g. \rightarrow Primary, secondary, Data, Questionnaire (ungrouped data) / Raw data \downarrow Data, Primary, Questionnaire, Secondary.

e.g. 13, 19, 12, 10, 25, 32, 29, 37 (ungrouped / Raw data) \downarrow Arrayed Data

  • 10, 12, 13, 19, 25, 29, 32, 37 \rightarrow Ascending / Increasing order
  • 37, 32, 29, 25, 19, 13, 12, 10 \rightarrow Descending / Decreasing order

5.1 Individual Series

Each item is given separate value. e.g

Name of studentWeight in (kgs)
X55
Y70
Z60
K55

5.2 Discrete Series (Discrete Frequency Distribution)

Each individual value is presented with frequency.

e.g.

Salary of EmployeesNo. of Employees(Frequency)
Tally mark
20,0005$\cancel{
40,0007$\cancel{
50,0004$
70,0002$
MarksNo. of studentsTally mark
13$
35$\cancel{
59$\cancel{
710$\cancel{
912$\cancel{

5.3 Continuous Series (Continuous Frequency Distribution)

Shows range of values of different items.

e.g.

Marks (class interval)No. of students
0-55
5-1010
10-1512
15-207
  • 0-5 \rightarrow class interval = 50=55 - 0 = \text{\textcircled{5}} (Range)
    • 00 \rightarrow lower limit
    • 55 \rightarrow upper limit
  • class mark = mid value = lower limit+upper limit2\frac{\text{lower limit} + \text{upper limit}}{2} =0+52=2.5 = \frac{0+5}{2} = \text{\textcircled{2.5}}

6. Tabulation of Data

🎯 THIS CAME IN THE EXAM — 2024 · Q74

"Which statement regarding tabulation is FALSE?" a) Tabulation allows presentation of complicated data · b) Tabulation is a prerequisite for diagrammatic representation · c) Statistical analysis necessarily involves tabulation · d) Tabulation facilitates comparison between rows and NOT columns ✅ ⭐ Answer: D — that is the FALSE one. A table compares across BOTH rows AND columns. The definition immediately below ("rows and columns") is exactly what kills it. ⚠️ Underline the word FALSE before you answer.

Systematic and Logical representation of numeric data in rows (\rightarrow horizontal) and columns (\downarrow vertical).

Objectives (1) To simplify the complex data. (2) To bringout important features. (3) To facilitate comparison. (4) To facilitate statistical analysis. (5) To save space and time.

Limitations (1) Lack of description (2) Incapable of presenting individual items. (3) Needs ample knowledge & understanding.

General Format

Table No: <Title> <Head note> (if any)

Stub
(Row heads)
Caption (column Headings)Total
(Rows)
Sub-Heads1Sub-Heads2Subheads1Sub-Heads2
<--BODY
-->
Total (cols)
  • S. Note: -
  • Foot note / Note

6.1 Main Parts of a Table

(1) Table NO

  • First item mentioned on top of table
  • Identification and References.

(2) Title \rightarrow

  • Second item, just above the table / by right side of T.NO.

(3) Head-note / Prefatory

  • 3rd item, just above the table.
  • Information about unit of data. e.g.: currency ₹ (or) $ Quintals or tonnes
  • Generally given in brackets.

(4) Caption / Col-Headings / characteristic

  • Top of each column.
  • Explains data in column.

(5) Stub / Row Heading

  • Title of horizontal rows.

(6) Body

  • Numeric data.
  • In cols and rows.

(7) Foot Note

  • To explain non-explanatory things.
  • Exceptions if any.
  • Circumstances affecting data.

(8) Source Note

  • Statement indicating source of data.

Table NO. 2.1 Records of Graduation students faculty wise H.Note: (Includes all registered students)

FacultyIst Year col. H2nd Year col. HTotal
(Row)
Boys S.HGirls S.HBoys S.HGirls S.H
B.Sc70305035185
B.com50203030130
B.A100606020240
Total (col)22011014085==555==

Source note: Admission dept. ABC college. Footnote: Late registrations are not included.

6.2 Characteristics of a Good Table

  • Title compatible with objective.
  • Comparable.
  • Ideal size.
  • Col-total, Row total, total included.
  • Stubs.
  • Headings.
  • Simple, Economical & Attractive.
  • No Abbreviations.
  • No Over crowding.

6.3 Types of Tables

graph TD
    Root[Types of Tables]
    
    Root --> Obj[objectivity / purpose]
    Root --> Nat[Nature / originality]
    Root --> Char[characteristic / constructio]
    
    Obj --> O1["==General Purpose=="]
    Obj --> O2["==Special Purpose=="]
    
    Nat --> N1["==Primary / original=="]
    Nat --> N2["==Secondary / Derived=="]
    
    Char --> C1["==Simple (1-way)=="]
    Char --> C2["==Complex=="]
    
    C2 --> C2_1["==Double / Two-way=="]
    C2 --> C2_2["==Treble / 3-way=="]
    C2 --> C2_3["==Manifold=="]

6.4 Classification of Tables

  • On the basis of "Purpose"General Purpose / Master Tables

    • General use.
    • Not meant for special purpose. e.g.: \rightarrow Census

    Special purpose table / Summary Table.

    • Derived from general table.
    • Serves special purpose.
    • Useful for calculation of analytical statistics like ratio, percentage etc. e.g.: \rightarrow Calculating Ratio of Male:Female
  • On the basis of "Originality" (1) Original Table / Primary Table

    • Data is presented in the form in which it is collected.

    (2) Derived Table / Secondary Table

    • Converted into any form as per requirement. e.g.: Limited columns as required.
  • As per construction / characteristics (1) Simple Table

    • "One-way" table.
    • Data presented based on the "one" characteristic only.

Table 1.1 Faculty-wise Number of students

Faculties: (Attribute)No. of Students
Science30
Commerce40
Arts60
Total130

Source note. One-way Table Foot note.

(2) Complex tables

  • More than one attribute presented simultaneously.

    (i) Double / Two-way Table

    • Data is tabulated on the basis of two inter-related characteristics.

Table 1.2 Faculty - wise number of Male and Female students

FACULTYNo. of Students (gender)Total
BOYSGIRLS
Arts352055
Commerce402565
Science303565
Total10580185

S. note. F. note.


  • Three-way / Trebles
    • Data is tabulated on the basis of 3-inter-related characteristics.

Table 1.3 Faculty wise, (Gender & semester) list of students

FACULTYNo. of studentsTotal
GirlsBoys
Sem ISem IITotal(1)Sem ISem IITotal(2)
Science152035205070105
Arts3530654585130195
Commerce2535603590125185
Total7585160100225325485

S. note. 3-way Table F. note.

  • Manifold (Higher order Table)
    • Data is tabulated on the basis of large no. of interrelated characteristics.

Table 1.4 Faculty wise (UG / PG) Gender based students in each sem.

FACULTYNumber of studentsTotal
UGPG
BoysGirlsBoysGirls
Sem IIITotalIIITotalIIITotalIIITotal
Total

S. note F. note


7. Graphical Representation of Data

Presentation of statistical data on graph paper in the form of lines or curves.

Merits (i) Simplifies complex data. (ii) Helps in forecasting. (iii) Variations in the values of variables.

Demerits (i) Precise values are not shown. (ii) May give wrong predictions. (iii) Difficult to interpret.

GraphVsDiagram
1. Graph paper is required.1. Can be drawn on plain paper.
2. Represents mathematical relationship.2. Does not represent any relationship.
3. Median, mode etc can be determined.3. Impossible through diagram.

7.1 Procedure to Construct a Graph

               y-axis (ordinate)
                      ^
                      |
         O₂           |           O₁
                      |
<---------------------+---------------------> x-axis
                      |                   (abscissa)
         O₃           |           O₄
                      |
                      v

(i) Scale Appropriate scale. (ii) Heading Suitable and precise heading. (iii) Proper Indications More than one curve, properly differentiated. (iv) False Base line y-axis. (v) Index lines & colors. (vi) Data Table full knowledge of data. (vii) Drawing line or curve Joining the points. (viii) Paper size Appropriate size of graph paper. (ix) From L to R & B to Top Left to right & Bottom to top.

7.2 Frequency Diagram (Histogram)

Frequency distribution is represented using graphs.

Histogram \rightarrow Two dimensional diagram, set of rectangles with base as the intervals between class boundaries. ** Histogram is not drawn for discrete data.

Equal classes

MarksNo. of Students
0 - 520
5 - 1030
10 - 1510
15 - 2040
20 - 2550
xychart-beta
title "Histogram (Student marks)"
x-axis ["0-5", "5-10", "10-15", "15-20", "20-25"]
y-axis "Students" 0 --> 50
bar [20, 30, 10, 40, 50]

Unequal classes

MarksNo of students
0 - 520
5 - 1530
15 - 2010
20 - 3040
30 - 3510
xychart-beta
title "Histogram (Student marks - Unequal classes)"
x-axis ["0-5", "5-15", "15-20", "20-30", "30-35"]
y-axis "Students" 0 --> 40
bar [20, 30, 10, 40, 10]

7.3 Bar Diagram

  • Bars with arbitrary width are drawn to represent the data.
  • Drawn for both discrete and continuous data.
  • Some space has to be left between consecutive bars.

e.g.: Production of Rice in J&K from 1997 - 2004

YearProduction (tonnes)
199710
199820
199930
200015
200120
200225
200330
200435
xychart-beta
title "Bar diagram: Production of Rice"
x-axis ["1997", "1998", "1999", "2000", "2001", "2002", "2003", "2004"]
y-axis "Production (tonnes)" 0 --> 40
bar [10, 20, 30, 15, 20, 25, 30, 35]

7.4 Frequency Polygon

Frequency polygon / Frequency polygon curve is constructed by joining the mid points of the frequencies presented using rectangles (histograms). ** constructed for both discrete as well as continuous series.

e.g.: Construct frequency polygon for continuous series

Age (Yrs)No. of students
5-105
10-1510
15-2015
20-2520
25-305
xychart-beta
title "Frequency Polygon for Age VS No. of students"
x-axis ["5-10", "10-15", "15-20", "20-25", "25-30"]
y-axis "No. of students" 0 --> 25
bar [5, 10, 15, 20, 5]
line [5, 10, 15, 20, 5]

e.g. construct Frequency polygon for discrete series

Age.No. of students
410
515
620
710
85
94
xychart-beta
title "Frequency polygon Age vs No. of students (discrete series)"
x-axis ["1", "2", "3", "4", "5", "6", "7", "8", "9", "10"]
y-axis "No. of students" 0 --> 25
line [0, 0, 0, 10, 15, 20, 10, 5, 4, 0]

7.5 Frequency Curve

Obtained by drawing smooth freehand curve passing through the closely associated points. ** Can be drawn for both discrete as well as continuous series.


➕ Added — How a Frequency Polygon BECOMES a Frequency Curve.

Histogram → Frequency Polygon → Frequency Curve is one continuous progression:

  1. Histogram — bars over each class
  2. Frequency polygon — join the mid-points of the tops of those bars
  3. Frequency curve — now make the class intervals SMALLER, so the number of classes INCREASES. Each corner of the polygon gets shallower, and the line becomes increasingly SMOOTH. In the limit (infinitely many, infinitely narrow classes) the polygon becomes the frequency curve.

[!pyq] 🎯 2022 · Q38"What happens to the shape of a frequency polygon as the number of classes increases?" → ⭐ It becomes increasingly smooth

Also worth knowing: the area under a frequency polygon equals the area of its histogram · a polygon can compare two or more distributions on one graph (a histogram cannot).


7.6 Types of Curves by Shape

  1. Normal curve: \rightarrow symmetric curve / Bell curve / Gaussian distribution
  2. U-shaped curve
  3. Positive skewed curve
  4. Negative skewed
  5. Bi-modal curve
  6. J-shaped curve
  7. Reverse J-shaped curve
  8. Mixed Multi-modal curve

7.7 Cumulative Frequency Curve (Ogive)

Salary (in Ks)No. of Employees (f)
10-207
20-3010
30-4015
40-5018
50-6020
60-7025

Less than (LTB) cumulative frequency \rightarrow Less than ogive curve:

Salaryless than cumulative
Less Than 207
3017
4032
5050
6070
7095

Greater than (GTB) cumulative frequency \rightarrow Greater than ogive curve:

SalaryMore than C.F.
1095
2088
3078
4063
5045
More than 6025

** intersection of Less than ogive & greater than ogive = Median

8. Class Intervals

8.1 True Class Intervals

There is no gap between the successive classes. Upper limit of each class is equal to lower limit of the succeeding class. e.g.

Age(No. of persons)
Frequency
10 - 207
20 - (30)8
(30) - 409
40 - 5010
  • Class Boundaries

8.2 Non-True Classes

Upper limit of each class is not equal to lower limit of successive class. e.g

Weight (kgs)No. of person
30 - (39)5
(40) - 4910
50 - 5920
60 - 6925
  • class limits

8.3 Open-End and Closed-End Classes

When a class limit is missing either at lower end (first class) or at upper end (last class) or limits are not specified at both ends — open end classes. When limits are specified at both ends — closed end classes.


Examples of open-end classes

MarksNo. of students
Below 2010
20 - 3012
30 - 4013
40 - 5015
MarksNo. of students
0 - 205
20 - 4010
40 - 6015
Above 6020
MarksNo. of students
Below 207
20 - 308
30 - 409
40 - 6020
Above 6012

Example of closed-end classes

MarksNo. of students
0 - 105
10 - 207
20 - 309
30 - 4012
40 - 5015

9. Important Points on Compilation and Tabulation

  1. Repository: Tables to show data in orderly manner.
  2. Marginal Frequency: Row totalGrand total=column totalGrand total\frac{\text{Row total}}{\text{Grand total}} = \frac{\text{column total}}{\text{Grand total}}
  3. Relative Frequency = Individual frequencyTotal no. of observations\frac{\text{Individual frequency}}{\text{Total no. of observations}}
  4. Univariate Frequency Distribution: One variable
  5. Bivariate Frequency Distribution: Two variables

END OF TOPIC (iv) — Topic (v) begins below.



&lt;- Previous · Index of all topics · Next ->

Built with LogoFlowershow