Statistics
Statistics
🗂️ THE 10 SYLLABUS TOPICS — how this file is divided
The FAA syllabus lists 10 statistics topics, and the 2024 paper asked exactly one question from each, in order. This file is being divided to match. Each topic starts with a blue TOPIC banner and ends with a grey END banner, so you always know which topic you are reading.
| # | Syllabus topic | Divided? |
|---|---|---|
| (i) | Primary and secondary data | ✅ |
| (ii) | Methods of collecting primary & secondary data | ✅ |
| (iii) | Preparation of questionnaires | ✅ |
| (iv) | Tabulation and compilation of data | ✅ |
| (v) | Measures of central tendency | ✅ |
| (vi) | Theory of probability | ✅ written in full — was entirely missing |
| (vii) | Theory of attributes | ✅ |
| (viii) | Theory of index numbers | ✅ |
| (ix) | Demography — Census | ✅ |
| (x) | Vital statistics | ✅ |
⬜ FOUNDATION — General Concepts
Background only — meaning, characteristics, limitations and basic vocabulary. Not one of the 10 syllabus topics. Topic (i) begins after this.
1. Five Basics / Fundamentals of Statistics
1.1 Variable
An attribute or characteristic of an item being analysed.
| Type | Number of variables | Example |
|---|---|---|
| Univariate | One | Height of a person in a country |
| Bivariate | Two | Height with respect to age |
| Multivariate | More than two | Height w.r.t. gender w.r.t. age |
1.2 Population / Universe
All the members, or the entire group of units, which is the focus of study. Example: all the citizens of a country; the residents of a particular geographical location.
1.3 Parameter
A characteristic used to define a given population — a numeric measure. Example: average income of every individual in a country.
1.4 Sample
A set of observations drawn from a population — a subset of the population selected for analysis. Example: a sample of a medicine.
1.5 Statistic
A characteristic that defines the given sample — a numerical measure of the sample. Example: number of RBCs in a blood sample.
The pairing to remember: a parameter describes a population; a statistic describes a sample.
2. Origin and Definition
Derivation of the word "Statistics":
graph TD
A[Word: Statistics] --> B["==Status==<br>(Latin word)"]
A --> C["==Statista==<br>(Italian word)"]
B --> D["Political State<br>(Political status)"]
C --> E["(Statesman)"]
Definition — Statistics is the study of:
graph TD
A["==Collection=="] --> B["==Classification / organisation=="]
B --> C["==Analysis=="]
C --> D["==Presentation / Interpretation=="]
D -.-> E[of Data]
Key people
| Title | Person |
|---|---|
| Founder of Statistics | John Graunt and William Petty (1662) |
| Father of Statistics | Sir Ronald Fisher (Fisher Principle / ANOVA) · Karl Pearson (graphical statistical techniques) |
| Father of Indian Statistics | Prasanta Chandra Mahalanobis (Mahalanobis distance · Kolkata Statistical Institute · Industrialisation Policy, 2nd FYP) |
What statistics is, and what it deals with
Statistics is a branch of Science / Mathematics / Mathematical Science.
graph TD
A[Data] --> B["Qualitative data<br>(Attributes / characteristics)"]
A --> C["Quantitative data<br>(Number / Quantity)"]
Data + Statistical Techniques = Statistics
Characteristics of statistics
- Data is numerically expressed
- A systematic manner is followed
- Expresses relations within the data (comparison)
- Provides metrical results
- Complexities in data are simplified
Fields of application
Economics · Commerce · Banking & Insurance · Public Opinion (exit polls, etc.)
Limitations — what statistics cannot do
- ❌ Study qualitative phenomena
- ❌ Reveal the entire background
- ❌ Be precise and exact (it works on averages)
- ❌ Resist misuse — the full context is always needed
3. Other Important Terms
| Term | Meaning |
|---|---|
| Generalizability | The ability to draw conclusions about the whole population from the results of a sample |
| Distribution | An arrangement of data by the values of one variable, in order from low to high |
| Skewness | The asymmetry of a distribution — see below |
Skewness and normal distribution
Height of Income
Students
. ˙ . . ˙ . . ˙ .
˙ ˙ Normal Bihar ˙ ˙ ˙ ˙ J&K
˙ ˙ Distribution ˙ ˙ ˙
+---+---+---+---+---+---+---+ +---+---+---+---+---+---+---+---+
0.8 1.0 1.2 1.4 1.6 1.8 2.0 5K 10K 15K 20K 25K 30K 40K
^ Left Tail Right Tail ^
(Negative) (Positive)
| Tail direction | Skew |
|---|---|
| Tail on the left | Negative skew |
| Tail on the right | Positive skew |
🔷 TOPIC (i) — PRIMARY AND SECONDARY DATA
Syllabus topic (i) of 10 · runs until Topic (ii)
Covers: data collection · primary data · secondary data · primary vs secondary · types of data · census vs sample and sampling error 🎯 Asked: 2024 · Q71 — secondary data is primarily sourced from?
1. Data Collection
The step after the research problem is identified and the research design is made is data collection.
Based on the approach to information gathering, data is categorised as:
graph TD
A[Data] --> B[Primary Data]
A --> C[Secondary Data]
2. Types of Data
➕ Added — not in the original notes.
2.1 By nature
| Quantitative (numerical) | Qualitative (categorical / attribute) |
|---|---|
| Measurable in numbers | Cannot be measured, only classified |
| Height, income, marks, age | Gender, literacy, blindness, religion |
| Handled by mean / median / mode | Handled by Theory of Attributes — topic (vii) |
2.2 Quantitative data splits again
| Variable | Takes | Example | Obtained by |
|---|---|---|---|
| Discrete | Only whole, separate values | Number of children, number of cars | Counting |
| Continuous | Any value in a range | Height, weight, temperature | Measuring |
2.3 By time reference
| Type | Meaning | Example |
|---|---|---|
| Time-series | One unit over many periods | India's GDP, 2010–2026 |
| Cross-sectional | Many units at one point in time | GDP of 20 countries in 2026 |
| Panel / longitudinal | Both together | GDP of 20 countries, 2010–2026 |
The quick test: count it in whole numbers → discrete · measure it to decimals → continuous · only label it → qualitative.
3. Primary Data
Data collected for the first time by the researcher himself, for a specific purpose. Example: a survey on job satisfaction of employees; health needs of a community.
| Advantages | Disadvantages |
|---|---|
| Investigator collects problem-specific data | More hectic and time consuming |
| No doubt about the quality of the data | Dealing with funding and funding agencies |
| Additional data can be obtained during the study | Ethical considerations (consent, permissions) |
| Costly / expensive | |
| Unnecessary or useless data is not included |
4. Secondary Data
a) Original research studies · b) Publicly available databases ✅ · c) Personal interviews · d) Laboratory experiments Answer: B. The three wrong options are all primary collection. Gather it yourself → primary; look it up → secondary.
Data that is already collected, produced or published by others. Example: use of hospital records or census records.
| Advantages | Disadvantages |
|---|---|
| Already collected — hassle free | Hard to get the specific data you need from it |
| Less expensive | Additional data or clarification cannot be obtained |
| Investigator is not responsible for data quality | Data may be of low quality or fabricated |
5. Primary vs Secondary Data
| # | Primary | Secondary |
|---|---|---|
| 1 | Real-time data | Past data |
| 2 | Time consuming | Quick and easy |
| 3 | Expensive | Economical |
| 4 | Available in crude form | Available in processed form |
| 5 | More accurate and reliable | Less accurate and reliable |
| 6 | Surveys, observations, experiments, questionnaires, schedules, local correspondents, interviews | Government publications, websites, books, journals, articles, internal records |
6. Census vs Sample, and Errors in Collection
➕ Added — not in the original notes.
6.1 Census (complete enumeration) vs Sample survey
| Census | Sample Survey | |
|---|---|---|
| Coverage | Every unit of the population | A part of the population |
| Cost / time | High | Low |
| Accuracy | High, if done well | Depends on the sample |
| Sampling error | NONE | Present |
| Non-sampling error | Present, often larger | Present, often smaller |
| Best when | Population is small; 100% detail needed | Population is large; units are destroyed on testing |
Testing that destroys the unit (bulbs, matchsticks, blood) must use sampling — a census would destroy the whole population.
6.2 Sampling error vs Non-sampling error
| Sampling Error | Non-Sampling Error | |
|---|---|---|
| Cause | Only a part of the population is studied | Mistakes in collection, recording, processing |
| Types | — | Measurement · coverage · non-response · response bias · processing |
| Occurs in a census? | No | Yes |
| As sample size increases | Decreases | May increase |
| Measurable? | Yes, statistically | Hard to measure |
a) It decreases as sample size increases ✅ The pair to remember: bigger sample → smaller sampling error, but possibly bigger non-sampling error.
6.3 Essentials of a good sample
Representative · Adequate (large enough) · Independent · Homogeneous · Free from bias
6.4 Sampling methods
| Random / Probability | Non-Random / Non-Probability |
|---|---|
| Simple random (lottery, random numbers) | Judgment / purposive |
| Stratified — split into homogeneous strata, sample each | Quota |
| Systematic — every k-th unit | Convenience |
| Cluster / multi-stage | Snowball |
In probability sampling every unit has a known, non-zero chance of selection — which is what makes the sampling error calculable.
🔷 TOPIC (ii) — METHODS OF COLLECTING PRIMARY AND SECONDARY DATA
Syllabus topic (ii) of 10 · runs until Topic (iv)
Covers: primary methods — experiment, interview, observation, questionnaire, schedules, local correspondents, projective techniques · secondary sources · precautions 🎯 Asked: 2024 · Q72 — example of primary data collection · 2022 · Q35 — method of collecting secondary data · 2022 · Q31 — sampling error
⚠️ Topic (iii) — Preparation of Questionnaires — is nested inside this topic, because the questionnaire is one of the primary collection methods. It is bracketed separately below.
PART A — PRIMARY DATA COLLECTION METHODS
Note: the original notes number these 1, 2 … then jump to 6–9 for the "other important methods" and 10 for projective techniques. That numbering is kept as written.
1. Experiment Method
Manipulating one variable to determine whether changes in it cause changes in another.
Three kinds of variables [x + y = z]
- Independent variable
- Dependent variable
- Controlled variable / controlled environment
| Adequacy — use when | Merits | Demerits |
|---|---|---|
| No outside factor may affect the outcome | Promises more accuracy | Results valid only in controlled conditions |
| A hypothesis must be validated | Reliable, bias-free data | Costly |
| Fair and unbiased data is required | Works with heterogeneous (varied) factors | Time consuming |
| Repeated trials are required | Uniquely isolates causal factors; highly controlled | Suits only simple, limited-scope problems |
| Scientific purposes |
2. Interview Method
a) Surveying existing research papers · b) Reviewing census data · c) Analyzing historical records · d) Conducting interviews and surveys ✅ Answer: D. The mirror image of Q71 — the paper tested the same primary/secondary split from both directions in one exam.
Based on oral or verbal stimuli.
graph TD
A[Interview Method] --> B[Direct Interview]
A --> C[Indirect Interview]
2.1 Direct Interview
The investigator directly contacts the respondent.
| Adequacy — use when | Merits | Demerits |
|---|---|---|
| The field of investigation is very small | Original data is collected | Costly and time consuming |
| A secret is to be kept | Accurate — data personally collected | Not suitable for a wide area |
| A high degree of originality is required | Comparative study is possible | Training and skill are required |
| Direct contact is required | Elastic — questions can be adjusted | Risk of investigator's personal bias |
2.2 Indirect Interview
Also called Indirect Personal Interview or Indirect Oral Interview. The investigator collects information from someone else who has the required information about the subject.
| Adequacy — use when | Merits | Demerits |
|---|---|---|
| The area of investigation is wide | Wide coverage | Less reliable — collected from a third party |
| Informants cannot be contacted directly | Independent of personal bias | Information providers may lack interest |
| Private information is being collected | Economical | Lack of uniformity |
| An expert point of view is needed |
2.3 Telephonic Interview
The dominant method. Questions are asked over the phone.
| Adequacy — use when | Merits | Demerits |
|---|---|---|
| The respondent is reluctant to be interviewed in person | Cheaper, and higher response than mailing | Reactions and gestures cannot be recorded (bias) |
| A personal interview is not possible | Time efficient | People without a phone connection cannot be reached |
| Wide coverage | Questions must be crisp and clear — difficult |
2.4 Other important interview methods
| Method | Meaning |
|---|---|
| (i) Structured | A set of predefined questions is used |
| (ii) Unstructured | No system of predetermined questions |
| (iii) Focussed | Attention focused on a given experience of the respondent and its possible effects |
| (iv) Clinical | Concerns broad underlying feelings and motivations, or the individual's life experience, rather than a specific experience |
| (v) Group | A group of 6 to 8 individuals is interviewed |
| (vi) Selection | For selecting people for certain jobs |
| (vii) Depth (in-depth) | Qualitative, unstructured, single respondent, highly skilled interviewer — to reach motivations, beliefs, attitudes and feelings |
2.5 Types of interview questions
| # | Type | Meaning |
|---|---|---|
| ① | Comprehension | Asks about prior experiences or employment |
| ② | Analytical | A problem with little information is given, to be solved |
| ③ | Open-ended / unrestricted | Cannot be answered with "Yes" or "No" |
| ④ | Closed-ended / restricted | Answerable in one word — "Yes", "No", etc. |
| ⑤ | Probing | Often begin with "What" or "How" to invite detail; "Do you"/"Are you" invite personal reflection |
| ⑥ | Contingency | Asked only if the respondent gives a particular response — follow-up questions |
| ⑦ | Leading | Push the respondent towards answering in a particular manner |
| ⑧ | Loading | Controversial questions containing presuppositions that force an answer |
⚠️ Leading and loading questions must be avoided in questionnaire design — see Topic (iii).
3. Observation Method
Data from the field is collected by the observer through observation.
| Adequacy — use when | Merits | Demerits |
|---|---|---|
| A hypothesis is to be tested | Current information is collected | Costly and time consuming |
| Non-verbal communication matters | Subjective bias is eradicated | Unforeseen factors may intrude |
| Independent of the respondent's manipulation | Respondent's opinion cannot be recorded on some subjects |
3.1 Classification of the observation method
graph TD
Root[Classification] --> Plan[Plan]
Root --> Conditions[Conditions]
Root --> Participation[Participation]
Plan --> Structured
Plan --> Unstructured
Conditions --> Controlled["Controlled<br>(Controlled conditions)"]
Conditions --> Uncontrolled["Uncontrolled<br>(Naturalistic conditions)"]
Participation --> C_Obs["==Complete observer=="]
Participation --> Obs_Part["==Observer as participant=="]
Participation --> Part_Obs["==Participant as observer=="]
Participation --> C_Part["==Complete participant=="]
C_Part -.-> Note["(Participant observation & Non-participant obs)"]
Disguised observation — the observer's identity is unknown to the subject. Example: "mystery shopping".
🔷 TOPIC (iii) — PREPARATION OF QUESTIONNAIRES
Syllabus topic (iii) of 10 · nested inside Topic (ii), because the questionnaire is a primary collection method
Covers: what a questionnaire is · merits and demerits · principles of preparation · construction steps (including the pre-test / pilot survey) · distribution · questionnaire vs schedule 🎯 Asked: 2024 · Q73 — which statements on questionnaire preparation are correct?
1. The Questionnaire
The heart of the survey operation.
A set of printed or written questions with a choice of answers, used for surveys or statistical studies — usually mailed.
Used to obtain: information · feelings · beliefs · perceptions about the research participant. Quantitative, qualitative and mixed data can all be collected through a questionnaire.
| Adequacy — use when | Merits | Demerits |
|---|---|---|
| Informants are educated | Less expensive | Only for educated respondents |
| The area of inquiry is wide | Free from the interviewer's influence | Possibility of misunderstanding |
| Information is supplied regularly | Respondents can take sufficient time to answer | Data received may not be reliable |
| Wide coverage | Lack of flexibility | |
| Less accuracy |
2. Principles for Preparing a Questionnaire
The TRUE statements: questionnaire design affects validity, reliability and response rate, and a pilot test / pre-test is essential. ⚠️ Any statement saying the pilot survey can be skipped is FALSE — see the construction flow below.
- Should be research oriented — the objective must be fulfilled
- The research participant should be kept in view
- Simple and familiar language should be used
- Questions should be clear, easy to understand and free of ambiguity
- Avoid leading and loading questions; avoid personal questions
- Avoid double-barrelled questions
- Avoid double-negative questions
- Use mutually exclusive and exhaustive response categories for closed-ended questions
- The questionnaire should be properly organised and easy to use
- Proper instructions for filling it must be provided
3. Construction of a Questionnaire
graph TD
S1["==Determine Research objective=="] --> S2["==Type of questionnaire to use=="]
S2 --> S3["==Determine the content of Individual=="]
S3 --> S4["==Type of Questions to use=="]
S4 --> S5["==Determine the wording of questions=="]
S5 --> S6["==Determining the sequence of Questions=="]
S6 --> S7["==Determining the length of Questionnaire=="]
S7 --> S8["==Layout=="]
S8 --> S9["==Check the Questions=="]
S9 --> S10["==Pre-test (Pilot survey)=="]
S10 --> S11["==Revision & Final draft=="]
⭐ Step 10 — the Pre-test (Pilot survey) — is the step the exam asks about. It comes after checking the questions and before the final draft.
4. Questionnaire Distribution
graph LR
A[Questionnaire distribution] --> B[Printed]
B --> B1[mailed]
B --> B2[Fax]
A --> C[Electronic]
C --> C1[email]
PART B — TOPIC (ii) RESUMED: REMAINING PRIMARY METHODS
Schedules, local correspondents and projective techniques are collection methods, so they belong to Topic (ii). The "Questionnaires vs Schedules" comparison below also relates to Topic (iii).
4. Schedules
A questionnaire filled in by enumerators. The investigator or enumerator personally visits informants with the questionnaire, asks the questions and notes the responses.
| Adequacy — use when | Merits | Demerits |
|---|---|---|
| Enough funds are available | Wide coverage | Expensive |
| Skilled, trained enumerators are available | Data is reliable | Time consuming |
| More accurate and reliable results are required | Fewer chances of bias | Trained enumerators are needed for reliable data |
| Direct contact with the respondent is possible | Can be used for illiterate respondents | Framing an ideal questionnaire is difficult |
| It is a questionnaire-cum-observation method |
5. Data from Local Correspondents
The investigator appoints a local correspondent or agent who collects information on the investigator's behalf.
| Adequacy — use when | Merits | Demerits |
|---|---|---|
| A regular supply of information is required | Wide coverage | Chances of personal bias |
| The area of investigation is wide | Cost efficient | Less accurate |
| A high degree of accuracy is not required | Time efficient | Lack of originality |
| Lack of uniformity |
6. Questionnaires vs Schedules
| # | Questionnaire | Schedule |
|---|---|---|
| 1 | Mailed | Filled by enumerator |
| 2 | Cheap | Expensive |
| 3 | Non-response high | Non-response low |
| 4 | Respondent not clearly known | Respondent known |
| 5 | Slow | Fast |
| 6 | Respondent must be literate | Not required |
| 7 | High risk of bias | Low risk of bias |
| 8 | Needs to be attractive | Not required — filled by the enumerator |
| 9 | Observation not possible | Observation can be used |
7–9. Other Important Methods
| # | Method | Meaning |
|---|---|---|
| 6 | Warranty cards | Postal-sized cards |
| 7 | Distributor / Store audits | Through distributors and manufacturers, via salesmen |
| 8 | Consumer panels | Consumers maintain detailed daily records |
| 9 | Mechanical / electronic devices | Cameras, ratings, reviews |
10. Projective Techniques
Also known as Indirect Interviewing.
- Used to reach underlying motives and intentions
- Applied where the respondent resists revealing them
- Requires intensive training
| # | Technique |
|---|---|
| (i) | Word association tests |
| (ii) | Sentence completion tests |
| (iii) | Story completion tests |
| (iv) | Quizzes |
| (v) | Verbal projection tests |
| (vi) | Pictorial techniques — see below |
10.1 Pictorial techniques
| Test | Key features |
|---|---|
| (a) Thematic Apperception Test (T.A.T.) | A set of pictures showing day-to-day or ambiguous events; responses are recorded |
| (b) Rosenzweig test | Cartoons with empty balloons inserted above |
| (c) Rorschach test | Ten cards with inkblots; symmetric but meaningless designs; used frequently, but validity is questioned |
| (d) Holtzman Inkblot Test (HIT) | A modification of the Rorschach; 45 cards; uses shading, movement and colour |
| (e) Tomkins-Horn Picture Arrangement Test | Designed for group administration; 25 plates, each with three sketches; the arrangement portrays a sequence of events |
PART C — SECONDARY DATA COLLECTION
Secondary data can be published or unpublished.
1. Published sources
a) Experiments · b) Personal Interview · c) Questionnaire · d) Government Publications ✅ Answer: D. Government publications are the classic secondary source, and the most-repeated example.
| # | Source | Examples |
|---|---|---|
| 1 | Government publications | Annual Survey of Industries · Agriculture Statistics Report · Indian Trade Journal |
| 2 | International organisations | UNO reports · WHO reports · World Bank Annual Report |
| 3 | Semi-official organisations | Reports of municipal corporations and district boards on health, sanitation, births, deaths |
| 4 | Committees and commissions | Bodies appointed by central or state governments — e.g. Tariff Commission, Patel Commission |
| 5 | Private publications | Journals · newspapers · research institutions · articles · reviews · reports |
2. Unpublished data
Some research institutes and universities do not publish their data, but it can still be used as a secondary source.
3. Precautions When Using Secondary Data
| Test | What to check |
|---|---|
| 1. Reliability | Source of the data · who collected it · when it was collected · methods used · bias in compilation · accuracy desired vs achieved |
| 2. Suitability | Is it suitable for the objective, nature and scope of the present enquiry? |
| 3. Adequacy | Is the level of accuracy achieved by this data adequate? |
Demerits of secondary data:
- Proper data-collection procedure may not have been used
- Out-dated data
- Bias
- Does not fulfil the accuracy needs of the research
- Not suitable for the period of investigation
4. One-Shot Recapitulation
| # | Primary | Secondary |
|---|---|---|
| 1 | Real-time data | Past data |
| 2 | Specific to the researcher's needs | Often not specific to the researcher's needs |
| 3 | Expensive | Less costly, or free |
| 4 | Time consuming | Time efficient |
| 5 | High control over quality | Less control over quality |
| 6 | Rudimentary form | Refined form |
| 7 | Collected by the original researcher | Collected by succeeding researchers |
| 8 | Researcher owns the data | Researcher may use it but cannot claim ownership |
Examples
| Primary | Secondary |
|---|---|
| Minutes of meetings, autobiographies | Fact books |
| Books and journal/newspaper articles written at the time of the event | Biographies |
| Documentaries, audio and video recordings | General histories |
| Maps, paintings, sculptures, drawings | Journals (other than those in the primary column) |
| Personal accounts | Books |
🔷 TOPIC (iv) — TABULATION AND COMPILATION OF DATA
Syllabus topic (iv) of 10 · runs until Topic (v)
Covers: Classification of data & its principles · statistical series · time / geographical / magnitude / condition classification · arrangement of data (individual, discrete, continuous) · tabulation & parts of a table · types of tables · graphical representation — histogram, bar diagram, frequency polygon, frequency curve, ogives · true, non-true & open-end classes 🎯 Asked in the papers: 2024 · Q74 — which statement about tabulation is FALSE · 2022 · Q38 — shape of a frequency polygon as classes increase
1. Classification of Data
Process of arranging data into different groups or classes based on common characteristics.
Characteristics of classification of Data
- It should be unambiguous
- Flexible to adjustments
- It should perform homogeneous grouping
- Its basis should be clearly defined and adhered to.
Objectives
- To simplify the huge data
- To facilitate comparison.
- To provide basis for tabulation
- To make data understandable.
- Clearly specify similarities and dissimilarities.
Representation of Data
graph LR
A[Representation of Data] --- B[Text]
A --- C[Tabular]
A --- D[Graphical]
2. Principles of Classification
**1. Exhaustive **
- Every item should be classified.
- NO Residual or miscellaneous class.
**2. Mutually Exclusive ** Every single item should fall in only one class.
**3. Stability ** Common principle should be maintained / followed for whole classification.
**4. Suitability ** Classification should be suitable as per objective of research.
**5. Flexibility ** Classification should be adjustable to new situation / new adjustments.
**6. Homogeneity ** Data items of a class should be similar in characteristics.
3. Statistical Series
Arrangement of data in a logical order.
graph TD
Root["==Statistical Series=="]
Root --> Char["On the basis of<br>characteristics"]
Root --> Const["On the basis of<br>construction"]
Char --> C1["Time-Series"]
Char --> C2["Spatial<br>series"]
Char --> C3["Condition<br>series"]
Char --> C4["Magnitude<br>series"]
C2 -.-> OR["(Or)"]
OR -.-> Simple["Simple<br>(one attribute)"]
OR -.-> Manifold["Manifold<br>(more than<br>one attributes)"]
Const --> S1["Individual<br>series"]
Const --> S2["Discrete<br>series"]
Const --> S3["Continuous<br>series"]
4. Important Classifications
4.1 Time-Series / Chronological / Temporal
"Demand: Sep 100, Oct 200, Nov 300, Dec 400. What is the 3-month moving average for January?" a) 300 ✅ · b) 350 · c) 400 · d) 450 ⭐ Answer: A. A 3-month moving average for January = mean of the previous three months = (200 + 300 + 400) ÷ 3 = 300. ⚠️ Moving averages are NOT in the 10-topic FAA statistics syllabus — this came from the 2022 Panchayat paper. Low priority, one line to learn.
Based on time of occurence. e.g population of India from 2010 - 2015
| Year | Population |
|---|---|
| 2010 | 106 cr. |
| 2011 | 110 cr. |
| 2012 | 114 cr. |
| 2013 | 120 cr. |
| 2014 | 126 cr. |
| 2015 | 130 cr. |
4.2 Geographical / Spatial
Classification on the basis of Area/Region/Location. e.g. sales report of company
| State | Sales (Lakhs) |
|---|---|
| Delhi | 30 |
| J&K | 20 |
| Punjab | 40 |
| Kolkata | 15 |
| Hyderabad | 30 |
4.3 Magnitude / Numerical / Quantitative
Based on quantity
| Height (ft) | No. of Persons |
|---|---|
| 5.10 | 25 |
| 5.11 | 15 |
| 6 | 10 |
| 6.1 | 5 |
4.4 Condition / Seasonal
When classification is done on the basis of seasonal variations/situations. e.g sales of ice cream / cold drinks in summers and winters.
**Simple classification ** Classification on the basis of only one attribute. e.g. On the basis of Gender
graph TD
A[Gender] --> B[Male]
A --> C[Female]
**Manifold classification ** Classification based on more than one attributes.
graph TD
Pop[Population] --> Lit[Literate]
Pop --> Illit[Illiterate]
Lit --> Emp1[Employed]
Lit --> Unemp1[Unemployed]
Illit --> Emp2[Employed]
Illit --> Unemp2[Unemployed]
Emp1 --> M1[Married]
Emp1 --> UM1[Unmarried]
Unemp1 --> M2[M]
Unemp1 --> UM2[UM]
Emp2 --> M3[M]
Emp2 --> UM3[UM]
Unemp2 --> M4[M]
Unemp2 --> UM4[UM]
5. Arrangement of Data
Data can be arranged in two ways:
- Serial (or) Alphabetic order
- Ascending (or) Descending order.
e.g. Primary, secondary, Data, Questionnaire (ungrouped data) / Raw data Data, Primary, Questionnaire, Secondary.
e.g. 13, 19, 12, 10, 25, 32, 29, 37 (ungrouped / Raw data)
Arrayed Data
10, 12, 13, 19, 25, 29, 32, 37Ascending / Increasing order37, 32, 29, 25, 19, 13, 12, 10Descending / Decreasing order
5.1 Individual Series
Each item is given separate value. e.g
| Name of student | Weight in (kgs) |
|---|---|
| X | 55 |
| Y | 70 |
| Z | 60 |
| K | 55 |
5.2 Discrete Series (Discrete Frequency Distribution)
Each individual value is presented with frequency.
e.g.
| Salary of Employees | No. of Employees | (Frequency) Tally mark |
|---|---|---|
| 20,000 | 5 | $\cancel{ |
| 40,000 | 7 | $\cancel{ |
| 50,000 | 4 | $ |
| 70,000 | 2 | $ |
| Marks | No. of students | Tally mark |
|---|---|---|
| 1 | 3 | $ |
| 3 | 5 | $\cancel{ |
| 5 | 9 | $\cancel{ |
| 7 | 10 | $\cancel{ |
| 9 | 12 | $\cancel{ |
5.3 Continuous Series (Continuous Frequency Distribution)
Shows range of values of different items.
e.g.
| Marks (class interval) | No. of students |
|---|---|
| 0-5 | 5 |
| 5-10 | 10 |
| 10-15 | 12 |
| 15-20 | 7 |
0-5class interval = (Range)- lower limit
- upper limit
- class mark = mid value =
6. Tabulation of Data
"Which statement regarding tabulation is FALSE?" a) Tabulation allows presentation of complicated data · b) Tabulation is a prerequisite for diagrammatic representation · c) Statistical analysis necessarily involves tabulation · d) Tabulation facilitates comparison between rows and NOT columns ✅ ⭐ Answer: D — that is the FALSE one. A table compares across BOTH rows AND columns. The definition immediately below ("rows and columns") is exactly what kills it. ⚠️ Underline the word FALSE before you answer.
Systematic and Logical representation of numeric data in rows ( horizontal) and columns ( vertical).
Objectives (1) To simplify the complex data. (2) To bringout important features. (3) To facilitate comparison. (4) To facilitate statistical analysis. (5) To save space and time.
Limitations (1) Lack of description (2) Incapable of presenting individual items. (3) Needs ample knowledge & understanding.
General Format
Table No: <Title>
<Head note> (if any)
| Stub (Row heads) | Caption (column Headings) | Total (Rows) | |||
|---|---|---|---|---|---|
| Sub-Heads1 | Sub-Heads2 | Subheads1 | Sub-Heads2 | ||
| <-- | BODY ↓ | --> | |||
| Total (cols) | |||||
- S. Note: -
- Foot note / Note
6.1 Main Parts of a Table
(1) Table NO
- First item mentioned on top of table
- Identification and References.
(2) Title
- Second item, just above the table / by right side of T.NO.
(3) Head-note / Prefatory
- 3rd item, just above the table.
- Information about unit of data. e.g.: currency ₹ (or) $ Quintals or tonnes
- Generally given in brackets.
(4) Caption / Col-Headings / characteristic
- Top of each column.
- Explains data in column.
(5) Stub / Row Heading
- Title of horizontal rows.
(6) Body
- Numeric data.
- In cols and rows.
(7) Foot Note
- To explain non-explanatory things.
- Exceptions if any.
- Circumstances affecting data.
(8) Source Note
- Statement indicating source of data.
Table NO. 2.1 Records of Graduation students faculty wise H.Note: (Includes all registered students)
| Faculty | Ist Year col. H | 2nd Year col. H | Total (Row) | ||
|---|---|---|---|---|---|
| Boys S.H | Girls S.H | Boys S.H | Girls S.H | ||
| B.Sc | 70 | 30 | 50 | 35 | 185 |
| B.com | 50 | 20 | 30 | 30 | 130 |
| B.A | 100 | 60 | 60 | 20 | 240 |
| Total (col) | 220 | 110 | 140 | 85 | ==555== |
Source note: Admission dept. ABC college. Footnote: Late registrations are not included.
6.2 Characteristics of a Good Table
- Title compatible with objective.
- Comparable.
- Ideal size.
- Col-total, Row total, total included.
- Stubs.
- Headings.
- Simple, Economical & Attractive.
- No Abbreviations.
- No Over crowding.
6.3 Types of Tables
graph TD
Root[Types of Tables]
Root --> Obj[objectivity / purpose]
Root --> Nat[Nature / originality]
Root --> Char[characteristic / constructio]
Obj --> O1["==General Purpose=="]
Obj --> O2["==Special Purpose=="]
Nat --> N1["==Primary / original=="]
Nat --> N2["==Secondary / Derived=="]
Char --> C1["==Simple (1-way)=="]
Char --> C2["==Complex=="]
C2 --> C2_1["==Double / Two-way=="]
C2 --> C2_2["==Treble / 3-way=="]
C2 --> C2_3["==Manifold=="]
6.4 Classification of Tables
-
On the basis of "Purpose" ① General Purpose / Master Tables
- General use.
- Not meant for special purpose. e.g.: Census
② Special purpose table / Summary Table.
- Derived from general table.
- Serves special purpose.
- Useful for calculation of analytical statistics like ratio, percentage etc. e.g.: Calculating Ratio of Male:Female
-
On the basis of "Originality" (1) Original Table / Primary Table
- Data is presented in the form in which it is collected.
(2) Derived Table / Secondary Table
- Converted into any form as per requirement. e.g.: Limited columns as required.
-
As per construction / characteristics (1) Simple Table
- "One-way" table.
- Data presented based on the "one" characteristic only.
Table 1.1 Faculty-wise Number of students
| Faculties: (Attribute) | No. of Students |
|---|---|
| Science | 30 |
| Commerce | 40 |
| Arts | 60 |
| Total | 130 |
Source note. One-way Table Foot note.
(2) Complex tables
-
More than one attribute presented simultaneously.
(i) Double / Two-way Table
- Data is tabulated on the basis of two inter-related characteristics.
Table 1.2 Faculty - wise number of Male and Female students
| FACULTY | No. of Students (gender) | Total | |
|---|---|---|---|
| BOYS | GIRLS | ||
| Arts | 35 | 20 | 55 |
| Commerce | 40 | 25 | 65 |
| Science | 30 | 35 | 65 |
| Total | 105 | 80 | 185 |
S. note. F. note.
- Three-way / Trebles
- Data is tabulated on the basis of 3-inter-related characteristics.
Table 1.3 Faculty wise, (Gender & semester) list of students
| FACULTY | No. of students | Total | |||||
|---|---|---|---|---|---|---|---|
| Girls | Boys | ||||||
| Sem I | Sem II | Total(1) | Sem I | Sem II | Total(2) | ||
| Science | 15 | 20 | 35 | 20 | 50 | 70 | 105 |
| Arts | 35 | 30 | 65 | 45 | 85 | 130 | 195 |
| Commerce | 25 | 35 | 60 | 35 | 90 | 125 | 185 |
| Total | 75 | 85 | 160 | 100 | 225 | 325 | 485 |
S. note. 3-way Table F. note.
- Manifold (Higher order Table)
- Data is tabulated on the basis of large no. of interrelated characteristics.
Table 1.4 Faculty wise (UG / PG) Gender based students in each sem.
| FACULTY | Number of students | Total | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| UG | PG | ||||||||||||
| Boys | Girls | Boys | Girls | ||||||||||
| Sem I | II | Total | I | II | Total | I | II | Total | I | II | Total | ||
| Total | |||||||||||||
S. note F. note
7. Graphical Representation of Data
Presentation of statistical data on graph paper in the form of lines or curves.
Merits (i) Simplifies complex data. (ii) Helps in forecasting. (iii) Variations in the values of variables.
Demerits (i) Precise values are not shown. (ii) May give wrong predictions. (iii) Difficult to interpret.
| Graph | Vs | Diagram |
|---|---|---|
| 1. Graph paper is required. | 1. Can be drawn on plain paper. | |
| 2. Represents mathematical relationship. | 2. Does not represent any relationship. | |
| 3. Median, mode etc can be determined. | 3. Impossible through diagram. |
7.1 Procedure to Construct a Graph
y-axis (ordinate)
^
|
O₂ | O₁
|
<---------------------+---------------------> x-axis
| (abscissa)
O₃ | O₄
|
v
(i) Scale Appropriate scale. (ii) Heading Suitable and precise heading. (iii) Proper Indications More than one curve, properly differentiated. (iv) False Base line y-axis. (v) Index lines & colors. (vi) Data Table full knowledge of data. (vii) Drawing line or curve Joining the points. (viii) Paper size Appropriate size of graph paper. (ix) From L to R & B to Top Left to right & Bottom to top.
7.2 Frequency Diagram (Histogram)
Frequency distribution is represented using graphs.
Histogram Two dimensional diagram, set of rectangles with base as the intervals between class boundaries. ** Histogram is not drawn for discrete data.
Equal classes
| Marks | No. of Students |
|---|---|
| 0 - 5 | 20 |
| 5 - 10 | 30 |
| 10 - 15 | 10 |
| 15 - 20 | 40 |
| 20 - 25 | 50 |
xychart-beta
title "Histogram (Student marks)"
x-axis ["0-5", "5-10", "10-15", "15-20", "20-25"]
y-axis "Students" 0 --> 50
bar [20, 30, 10, 40, 50]
Unequal classes
| Marks | No of students |
|---|---|
| 0 - 5 | 20 |
| 5 - 15 | 30 |
| 15 - 20 | 10 |
| 20 - 30 | 40 |
| 30 - 35 | 10 |
xychart-beta
title "Histogram (Student marks - Unequal classes)"
x-axis ["0-5", "5-15", "15-20", "20-30", "30-35"]
y-axis "Students" 0 --> 40
bar [20, 30, 10, 40, 10]
7.3 Bar Diagram
- Bars with arbitrary width are drawn to represent the data.
- Drawn for both discrete and continuous data.
- Some space has to be left between consecutive bars.
e.g.: Production of Rice in J&K from 1997 - 2004
| Year | Production (tonnes) |
|---|---|
| 1997 | 10 |
| 1998 | 20 |
| 1999 | 30 |
| 2000 | 15 |
| 2001 | 20 |
| 2002 | 25 |
| 2003 | 30 |
| 2004 | 35 |
xychart-beta
title "Bar diagram: Production of Rice"
x-axis ["1997", "1998", "1999", "2000", "2001", "2002", "2003", "2004"]
y-axis "Production (tonnes)" 0 --> 40
bar [10, 20, 30, 15, 20, 25, 30, 35]
7.4 Frequency Polygon
Frequency polygon / Frequency polygon curve is constructed by joining the mid points of the frequencies presented using rectangles (histograms). ** constructed for both discrete as well as continuous series.
e.g.: Construct frequency polygon for continuous series
| Age (Yrs) | No. of students |
|---|---|
| 5-10 | 5 |
| 10-15 | 10 |
| 15-20 | 15 |
| 20-25 | 20 |
| 25-30 | 5 |
xychart-beta
title "Frequency Polygon for Age VS No. of students"
x-axis ["5-10", "10-15", "15-20", "20-25", "25-30"]
y-axis "No. of students" 0 --> 25
bar [5, 10, 15, 20, 5]
line [5, 10, 15, 20, 5]
e.g. construct Frequency polygon for discrete series
| Age. | No. of students |
|---|---|
| 4 | 10 |
| 5 | 15 |
| 6 | 20 |
| 7 | 10 |
| 8 | 5 |
| 9 | 4 |
xychart-beta
title "Frequency polygon Age vs No. of students (discrete series)"
x-axis ["1", "2", "3", "4", "5", "6", "7", "8", "9", "10"]
y-axis "No. of students" 0 --> 25
line [0, 0, 0, 10, 15, 20, 10, 5, 4, 0]
7.5 Frequency Curve
Obtained by drawing smooth freehand curve passing through the closely associated points. ** Can be drawn for both discrete as well as continuous series.
➕ Added — How a Frequency Polygon BECOMES a Frequency Curve.
Histogram → Frequency Polygon → Frequency Curve is one continuous progression:
- Histogram — bars over each class
- Frequency polygon — join the mid-points of the tops of those bars
- ⭐ Frequency curve — now make the class intervals SMALLER, so the number of classes INCREASES. Each corner of the polygon gets shallower, and the line becomes increasingly SMOOTH. In the limit (infinitely many, infinitely narrow classes) the polygon becomes the frequency curve.
[!pyq] 🎯 2022 · Q38 — "What happens to the shape of a frequency polygon as the number of classes increases?" → ⭐ It becomes increasingly smooth
Also worth knowing: the area under a frequency polygon equals the area of its histogram · a polygon can compare two or more distributions on one graph (a histogram cannot).
7.6 Types of Curves by Shape
- Normal curve: symmetric curve / Bell curve / Gaussian distribution
- U-shaped curve
- Positive skewed curve
- Negative skewed
- Bi-modal curve
- J-shaped curve
- Reverse J-shaped curve
- Mixed Multi-modal curve
7.7 Cumulative Frequency Curve (Ogive)
| Salary (in Ks) | No. of Employees (f) |
|---|---|
| 10-20 | 7 |
| 20-30 | 10 |
| 30-40 | 15 |
| 40-50 | 18 |
| 50-60 | 20 |
| 60-70 | 25 |
Less than (LTB) cumulative frequency Less than ogive curve:
| Salary | less than cumulative |
|---|---|
| Less Than 20 | 7 |
| 30 | 17 |
| 40 | 32 |
| 50 | 50 |
| 60 | 70 |
| 70 | 95 |
Greater than (GTB) cumulative frequency Greater than ogive curve:
| Salary | More than C.F. |
|---|---|
| 10 | 95 |
| 20 | 88 |
| 30 | 78 |
| 40 | 63 |
| 50 | 45 |
| More than 60 | 25 |
** intersection of Less than ogive & greater than ogive = Median
8. Class Intervals
8.1 True Class Intervals
There is no gap between the successive classes. Upper limit of each class is equal to lower limit of the succeeding class. e.g.
| Age | (No. of persons) Frequency |
|---|---|
| 10 - 20 | 7 |
| 20 - (30) | 8 |
| (30) - 40 | 9 |
| 40 - 50 | 10 |
- Class Boundaries
8.2 Non-True Classes
Upper limit of each class is not equal to lower limit of successive class. e.g
| Weight (kgs) | No. of person |
|---|---|
| 30 - (39) | 5 |
| (40) - 49 | 10 |
| 50 - 59 | 20 |
| 60 - 69 | 25 |
- class limits
8.3 Open-End and Closed-End Classes
When a class limit is missing either at lower end (first class) or at upper end (last class) or limits are not specified at both ends — open end classes. When limits are specified at both ends — closed end classes.
Examples of open-end classes
| Marks | No. of students |
|---|---|
| Below 20 | 10 |
| 20 - 30 | 12 |
| 30 - 40 | 13 |
| 40 - 50 | 15 |
| Marks | No. of students |
|---|---|
| 0 - 20 | 5 |
| 20 - 40 | 10 |
| 40 - 60 | 15 |
| Above 60 | 20 |
| Marks | No. of students |
|---|---|
| Below 20 | 7 |
| 20 - 30 | 8 |
| 30 - 40 | 9 |
| 40 - 60 | 20 |
| Above 60 | 12 |
Example of closed-end classes
| Marks | No. of students |
|---|---|
| 0 - 10 | 5 |
| 10 - 20 | 7 |
| 20 - 30 | 9 |
| 30 - 40 | 12 |
| 40 - 50 | 15 |
9. Important Points on Compilation and Tabulation
- Repository: Tables to show data in orderly manner.
- Marginal Frequency:
- Relative Frequency =
- Univariate Frequency Distribution: One variable
- Bivariate Frequency Distribution: Two variables
🔷 TOPIC (v) — MEASURES OF CENTRAL TENDENCY
Syllabus topic (v) of 10 · runs until Topic (vi)
Covers: Construction of frequency distributions · inclusive vs exclusive method · Mean (individual, discrete, continuous; combined mean, missing values, correction, weighted mean) · Median · Mode · ➕ added: empirical relation, comparison of the three averages, GM & HM, partition values 🎯 Asked in the papers: 2024 · Q75 — arrange mean / median / mode / n in decreasing order
1. Construction of a Frequency Distribution
Unarranged data can be distributed into classes and corresponding frequencies be assigned. Steps to follow
- Sorting of Data.
- Calculate the Range.
- Decide number of classes.
- Calculate class width.
- Tally & count the observations.
e.g. 0, 1, 4, 3, 12, 13, 19, 15, 20, 23, 27, 29, 37, 31, 41, 46, 50, 48 18 observations.
Step 1 — Sorting the data
Step 2 — calculate the range { 0 - 50 } = 50
Step 3 — Decide the number of classes. Let it be '5'
Step 4 — calculate class width.
| class | frequency |
|---|---|
| 0 - 10 | 4 |
| 10 - 20 | 4 |
| 20 - 30 | 4 |
| 30 - 40 | 2 |
| 40 - 50 | 4 |
Steps count observations & tally. No. of observations = frequency.
1.1 Inclusive vs Exclusive Method
| Wt. (Kg) | No. of student |
|---|---|
| 20 - 29 | 5 |
| 30 - 39 | 10 |
| 40 - 49 | 15 |
| 50 - 59 | 12 |
- Inclusive Method / series (Both limits are included)
conversion to Exclusive Method distribution
| Wt (Kg.) | No. of students |
|---|---|
| [(20 - 0.5) - (29 + 0.5)] 19.5 - 29.5 | 5 |
| 29.5 - 39.5 | 10 |
| 39.5 - 49.5 | 15 |
| 49.5 - 59.5 | 12 |
2. Mean
Definition — the average, or most common value, in a collection of numbers.
Two kinds:
- Simple Arithmetic Mean
- Weighted Mean
2.1 Simple Arithmetic Mean
Average of the collection of numbers. Formulas for calculating Simple Arithmetic Mean:
| TYPE OF SERIES | Direct Method | Assumed / Indirect Method |
|---|---|---|
| 1. Individual Series | ||
| 2. Discrete Series | ||
| 3. Continuous Series |
2.2 Individual Series
Q. Find the A. Mean of the given series:
Sol
| S.No | Marks (x) |
|---|---|
| 1 | 5 |
| 2 | 10 |
| 3 | 15 |
| 4 | 20 |
| 5 | 25 |
2.3 Discrete Series
| x | f | fx |
|---|---|---|
| 3 | 5 | 15 |
| 4 | 2 | 8 |
| 5 | 1 | 5 |
| 6 | 3 | 18 |
| 7 | 6 | 42 |
Note: N =
Now
2.4 Continuous Series
| classes | f | fx | |
|---|---|---|---|
| 0 - 10 | 2 | 5 | 10 |
| 10 - 20 | 3 | 15 | 45 |
| 20 - 30 | 4 | 25 | 100 |
| 30 - 40 | 2 | 35 | 70 |
| 40 - 50 | 5 | 45 | 225 |
| 50 - 60 | 6 | 55 | 330 |
Now
2.5 Combined Mean
| (Average Income) | (f) No. of persons |
|---|---|
2.6 Finding a Missing Value
Q. The mean temperature for four days noted is 120°c. If the temperature for day 1, day 2 & day 4 is 30, 35 and 40 respectively. Find the temperature of day 3?
Sol
Q. Find the missing value if the mean value = 10.
| x | f | fx |
|---|---|---|
| 2 | 5 | 10 |
| 3 | 10 | 30 |
| 4 | 15 | 60 |
| 20 | ||
| 6 | 25 | 150 |
2.7 Correcting a Mean Value
Q. In a class, average marks of 30 students is 40. If the correct marks for one student is 46 which was misread as 42. calculate the new correct mean.
Sol
Sum of all the marks.
- Add correct marks
- Subtract wrong marks.
New . New correct mean
2.8 Weighted Arithmetic Mean
When the items of a series are not of equal importance / weightage. where weighted Mean. sum of product of weights and items. sum of weights.
Q. student scores 20 marks in statistics, 35 in english, 40 in maths and 45 in Geography. calculate weighted mean, if the marks are weighted as 2, 1, 3, 4 respectively.
Sol
| marks (x) | weights (w) | wx |
|---|---|---|
| 20 | 2 | 40 |
| 35 | 1 | 35 |
| 40 | 3 | 120 |
| 45 | 4 | 180 |
3. Median
Definition — the middle value, which separates the higher half from the lower half of the data.
3.1 Individual Series
- Sort the data in ascending or descending order
- Calculate the median by the formula:
e.g 1 Sol Sort: Trick: (odd) Middle value = Median. here
e.g 2 Sol Already sorted Trick:
e.g 3
- sort values of X.
| S.No | x |
|---|---|
| 1 | 5 |
| 2 | 10 |
| 3 | 15 |
| 4 | (20) |
| 5 | 25 |
| 6 | 30 |
| 7 | 35 |
3.2 Discrete Series
(X) (frequencies) *1. Sort values in Ascending or decending order. *2. Cummulate frequencies *3.
e.g
| X | F | C.F |
|---|---|---|
| 10 | 5 | 5 |
| 20 | 6 | 11 |
| 30 | 7 | 18 |
| (40) | 10 | (28) |
| 50 | 15 | 43 |
| 60 | 20 | 63 |
check class and corresponding X-value
3.3 Continuous Series
(1) Sort groups in ascending or decending order. (2) change inclusive series into exclusive series. (3) cummulate frequencies (4) Calculate
e.g
| Marks | F | C.F |
|---|---|---|
| 0 - 10 | 5 | 5 |
| 10 - 20 | 3 | 8 |
| 20 - 30 | 2 | 10 |
| 30 - 40 | 7 | 17 |
| 40 - 50 | 6 | 23 |
| 80 - 100 | 4 | 27 |
Now
4. Mode (Z)
Definition — the value in the data set that appears most frequently; the value repeated the maximum number of times.
4.1 Individual Series
observation that is repeated maximum times e.g 1, 2, 3, 4, , 6, 5, 9, 10, 6, , 1 Mode = 3 (repeated maximum times)
e.g 2 1, 2, 3, 6, 7, 9 All Mode (no observation repeated).
e.g 3 4, , 4, 2, , 5, 6, 8. Mode = 3 & 4 (Bimodal).
- If there are more than two modes in a data set, it is called Multi-modal data.
[!fix] ❗ MISSING FORMULA — the empirical relation These notes cover Mean, Median and Mode thoroughly but never state the relation that links them:
⭐ Mode = 3 Median − 2 Mean
Equivalently Mean − Mode = 3 (Mean − Median). For a symmetric distribution Mean = Median = Mode. This is the standard one-mark question of this topic and is needed for 2024·Q75-type items where all three must be ordered.
4.2 Discrete Series
Observation that is corresponding to the maximum frequency.
| Age(x) | No. of persons(f) |
|---|---|
| 10 | 2 |
| 12 | 4 |
| 15 | 8 |
| 20 | (10) highest / maximum frequency |
| 25 | 3 |
| 30 | 5 |
4.3 Continuous Series
| Marks | No. of Students (f) |
|---|---|
| 10 - 20 | 5 |
| 20 - 30 | 8 |
| 30 - 40 | 10 |
| 40 - 50 | 12 |
| 50 - 60 | 6 |
| 60 - 70 | 3 |
| 70 - 80 | 2 |
Sol
- check if for exclusiveness
- select the maximum/highest frequency
- its corresponding class becomes modal class. {40-50}
Now
If in the above example mean is increased by 2, what will happen to the individual observation if all are equally affected.
Sol
If mean is increased by 2 then New mean =
Now
This 10 needs to be equally distributed. each observation will get increment.
➕ Added — Relations Between the Averages, and Partition Values.
5. Relations Between the Averages
5.1 The Empirical Relation
Rearranged forms you may need:
- Also written Mean − Mode = 3 (Mean − Median)
- ⭐ For a SYMMETRIC distribution: Mean = Median = Mode
- Positively skewed: Mean > Median > Mode · Negatively skewed: Mean < Median < Mode
e.g. Mean = 25, Median = 24 → Mode = 3(24) − 2(25) = 72 − 50 = 22
5.2 Comparison of the Three Averages
"Eight students' sleeping hours: 4, 8, 7, 5, 3, 7, 7, 3. Arrange in DECREASING order: a. Mean · b. Median · c. Mode · d. Number of sample" a) a,b,c,d · b) a,d,b,c · c) d, c, b, a ✅ · d) c,b,d,a ⭐ Working: sorted → 3,3,4,5,7,7,7,8. n = 8 · Mode = 7 · Median = (5+7)/2 = 6 · Mean = 44/8 = 5.5 So 8 > 7 > 6 > 5.5 = d, c, b, a. ⚠️ Note the trick: "number of sample" (n) is one of the four items — it is not a measure of central tendency at all, and it is the largest.
| Mean | Median | Mode | |
|---|---|---|---|
| Type | Mathematical average | Positional average | Positional average |
| Uses all observations? | ✅ Yes | ❌ No | ❌ No |
| ⭐ Affected by extreme values? | ⭐ YES | ⭐ NO | ⭐ NO |
| Can be found with open-end classes? | ❌ No | ✅ Yes | ✅ Yes |
| Can be located graphically? | ❌ No | ✅ Ogive | ✅ Histogram |
| Can there be more than one? | No | No | ⭐ Yes (bi-/multi-modal) — or none |
| Suitable for qualitative data? | No | Yes | Yes |
| Sum of deviations from it = 0? | ⭐ Yes (Σ(x−x̄)=0) | No | No |
⭐ Σ(x − x̄)² is MINIMUM when taken about the MEAN.
5.3 Types of Averages
Mathematical: Arithmetic Mean (AM) · Geometric Mean (GM) · Harmonic Mean (HM) Positional: Median · Mode
⭐ AM ≥ GM ≥ HM (always; equal only when all values are identical) · ⭐ GM² = AM × HM Use GM for rates of growth / ratios / index numbers · Use HM for speeds and rates per unit.
6. Partition Values — Quartiles, Deciles, Percentiles
Values that divide an ordered data set into equal parts.
| Measure | Divides into | Count | Middle value |
|---|---|---|---|
| Quartiles | 4 parts | Q₁, Q₂, Q₃ | ⭐ Q₂ = Median |
| Deciles | 10 parts | D₁ … D₉ | ⭐ D₅ = Median |
| Percentiles | 100 parts | P₁ … P₉₉ | ⭐ P₅₀ = Median |
⭐ Also: Q₁ = P₂₅ · Q₃ = P₇₅ · D₁ = P₁₀
Individual & discrete series (N = number of items):
Continuous series (interpolation, cf = cumulative frequency of the preceding class):
Related: Quartile Deviation (Semi-Interquartile Range) · Interquartile Range
🔷 TOPIC (vi) — THEORY OF PROBABILITY
Syllabus topic (vi) of 10 · runs until Topic (vii)
Covers: basic terms · types of events · the three definitions of probability · addition theorem · multiplication theorem · conditional probability · odds · Bayes theorem 🎯 Asked in the papers: 2024 · Q76 — which probability statements are correct · 2022 · Q34 — P(Maths or Physics)
➕ Added — not in the original notes.
1. Basic Terms
| Term | Meaning |
|---|---|
| Random experiment | An act whose outcome cannot be predicted with certainty (tossing a coin, rolling a die) |
| Trial | One performance of the experiment |
| Outcome | A single possible result |
| Sample space (S) | The set of ALL possible outcomes. Die: , so |
| Event (A) | Any subset of the sample space |
| Favourable outcomes | Outcomes that make the event happen |
2. Types of Events
"Which statements about probability are correct?" → Answer: A (P and S) ⭐ This question is decided entirely by the definitions in the table below. The two traps it used:
- "Mutually exclusive events always sum to 1" → ❌ FALSE — only if they are ALSO EXHAUSTIVE
- "Probability can never be zero" → ❌ FALSE — an impossible event is exactly 0 ⚠️ Mutually exclusive · exhaustive · independent are three different things. The paper mixes them on purpose.
| Type | Meaning | Example |
|---|---|---|
| Simple / Elementary | A single outcome | Getting a 4 on a die |
| Compound | More than one outcome | Getting an even number |
| Sure / Certain | Always happens → ⭐ P = 1 | A number less than 7 on a die |
| Impossible | Can never happen → ⭐ P = 0 | Getting 8 on a die |
| ⭐ Mutually exclusive | Cannot happen together → | Head and Tail on one toss |
| ⭐ Exhaustive | Together cover the whole sample space → total probability = 1 | {even, odd} on a die |
| ⭐ Independent | One does not affect the other | Two separate coin tosses |
| Dependent | One does affect the other | Drawing 2 cards without replacement |
| Complementary () | "A does not happen" → ⭐ | Not getting a six |
| Equally likely | All outcomes have the same chance | A fair die |
3. Definitions of Probability
(1) Classical / Mathematical (a priori): Requires outcomes to be equally likely, mutually exclusive and exhaustive.
(2) Empirical / Statistical (a posteriori): based on actual repeated trials —
(3) Axiomatic (Kolmogorov): ⭐ · · for mutually exclusive events
Impossible event = 0 · Certain event = 1 · Probability can NEVER be negative and NEVER exceed 1. ⭐ In the exam, any option greater than 1 (or negative) is instantly wrong.
4. Addition Theorem — "OR" / union
"80 students: 30 opted Maths, 20 opted Physics, 10 opted both. Find P(Maths or Physics)." a) 1/2 ✅ · b) 1½ · c) 2½ · d) 3½ ⭐ Working: P(M) = 30/80, P(P) = 20/80, P(M∩P) = 10/80 P(M ∪ P) = 30/80 + 20/80 − 10/80 = 40/80 = 1/2 ⚠️ ⭐ Look at options b, c and d — every one of them is GREATER THAN 1, so none can be a probability. The range rule alone eliminates three of the four options before you calculate anything.
General (works always):
⭐ If A and B are mutually exclusive, , so:
Three events:
5. Multiplication Theorem — "AND" / intersection
⭐ If A and B are INDEPENDENT:
If DEPENDENT:
6. Conditional Probability
⭐ If A and B are independent, — knowing B tells you nothing about A.
7. Odds
- Odds in favour of A = = favourable : unfavourable
- Odds against A =
- If odds in favour are then
8. Bayes Theorem
Used to revise a prior probability after new evidence arrives.
9. Worked Examples
Q1. (the 2022 · Q34 type) In a class, P(passing Maths) = 2/5, P(passing Physics) = 3/10, P(passing both) = 1/5. Find P(passing Maths or Physics).
Sol Use the general addition theorem —
Q2. A die is thrown once. Find the probability of getting an even number or a number greater than 4.
Sol ; ;
Q3. Two coins are tossed. Find P(at least one head).
Sol , . Easier by complement — ⭐ "At least one" is almost always fastest via the complement.
Q4. A bag has 5 red and 3 black balls. Two are drawn with replacement. P(both red)?
Sol With replacement → independent — Without replacement (dependent):
| Statement | Verdict |
|---|---|
| "Probability can never be zero" | ❌ FALSE — an impossible event is exactly 0 |
| "Probability of a certain event is 1" | ✅ TRUE |
| "Mutually exclusive events always sum to 1" | ❌ FALSE — ⭐ only if they are ALSO EXHAUSTIVE |
| "Mutually exclusive means independent" | ❌ FALSE — opposites in effect: if A happens B cannot, so they are strongly dependent |
| "P(A) + P(not A) = 1" | ✅ TRUE |
| "Probability can exceed 1 if there are many outcomes" | ❌ FALSE — never |
| "For independent events P(A and B) = P(A) × P(B)" | ✅ TRUE |
| "P(A or B) = P(A) + P(B) always" | ❌ FALSE — only when mutually exclusive; otherwise subtract |
⭐ — any option above 1 is eliminable free (this killed 3 of 4 options in 2022 · Q34) ⭐ OR → add, then subtract the overlap · AND → multiply ⭐ Mutually exclusive ≠ exhaustive ≠ independent — the three words the paper mixes up on purpose
🔷 TOPIC (vii) — THEORY OF ATTRIBUTES
Syllabus topic (vii) of 10 · runs until Topic (viii)
Covers: Attributes & notation · number of classes · order of frequencies · algebraic expressions · contingency tables · consistency of data · independence & association — proportion method and (AB) = (A)(B)/N · Yule's coefficient of association 🎯 Asked in the papers: 2024 · Q77 — class of 100 students, match the boys/girls figures · 2022 · Q39 — ultimate class frequency for independent attributes
Theory of Attributes
Deals with the qualitative characteristics calculated using quantitative measurements. e.g:- honesty, habit of smoking etc.
Attributes:- "Qualitative characteristics of an individual."
Dicotomony / Dichotomous classification:- Attribute divides class into two at each level.
graph TD
A[Population] --> B[Male]
A --> C[Female]
B --> D[Literate]
B --> E[Illiterate]
C --> F[Literate]
C --> G[Illiterate]
Types of Attributes
graph TD
A[Types of Attributes] --> B["Positive Attributes<br>(Presence)<br>Attribute"]
A --> C["Negative Attributes<br>(Absence)<br>Attribute"]
Symbols / Notations / Mnemonics:- Presence of Attribute:- Capital Letters; Absence of Attribute:- Greek Letters;
e.g:-
- class:- Homogeneous & Mutually exclusive group.
- class Frequency:- Number of items in each class.
- denoted by brackets over class symbols. e.g:
Number of classes:-
1 - Attribute — 3 classes — 2 - Attribute — 9 classes . 3 - Attributes — 27 classes. n - Attributes =
e.g No. of Attributes = 7.
Order of frequencies:-
Order - 0 : Order - 1 : Order - 2 : Order - 3 :
Some Algebraic Expressions:-
Formula for any order frequency classes:-
No. of classes of order frequency
- No. of attributes
- order
Q,, Find the number of Order-2 classes if the number of attributes is 2. Sol:- Number of 2 order class =
Q,, Calculate the 1st-order classes if the number of attributes = 2. Sol:- No. of Order-1 classes =
Q,, calculate No. of order-2 classes if the number of attributes = 3 Sol:- No. of order-2 classes =
Contingency table:
100 students, one sport each: Football 28, Kho-Kho 27, Volleyball 33, Cricket 12. 23 boys play football, 13 boys play volleyball. Of a total 40 girls, 13 play Kho-Kho. Match the figures. → Answer: D (a-3, b-1, c-2, d-5) ⭐ Method — build the 2-way table and fill the gaps:
- Total 100, girls 40 → ⭐ boys = 60
- Kho-Kho 27, girls 13 → boys Kho-Kho = 14
- Boys so far: football 23 + volleyball 13 + kho-kho 14 = 50 → ⭐ boys cricket = 60 − 50 = 10
- Cricket 12 total → girls cricket = 2 ⚠️ Source note: the paper prints "23 boys play cricket", which cannot be true (cricket has only 12 players in total). It must read football — with that reading the 60/40/100 grid balances perfectly. Treat it as an OCR/typo error in the paper. Matrix for a frequency table.
| Attributes | A | Total | |
|---|---|---|---|
| B | (AB) | (B) | (B) |
| (A) | () | () | |
| Total | (A) | () | N Population |
Note: Ultimate frequencies Highest order Frequencies
- Contingency matrix shows the relation between two variables / attributes.
- Used by Karl Pearson for first time — theory of contingency Relation to Association & Normal correlation
Contingency table / Cross-Tabulation / Cross-Tab.
graph TD
A[Contingency table] --> B["==2-Attributes=="]
A --> C[3-Attributes]
B -.-> B1[* 9 square table]
B -.-> B2[* 2x2 table]
Practice Questions on Contingency matrix
Q1. If (A) = 40, (AB) = 50, B = 80 & N = 160. Find out the remaining values. Sol:-
| Attributes | A | Total | |
|---|---|---|---|
| B | (AB) 50 | (B) 30 | (B) 80 |
| (A) -10 | () 90 | () 80 | |
| Total | (A) 40 | () 120 | N 160 |
Q2:- If the values (AB) = 60, (A) = 40, (B) = 30, () = 20. Find out the rest of values. Sol:-
| Attributes | A | Total | |
|---|---|---|---|
| B | (AB) 60 | (B) 30 | (B) 90 |
| (A) 40 | () 20 | () 60 | |
| Total | (A) 100 | () 50 | N 150 |
H/W:- Q3: If (A)=130, (B)=100, ()=120 and ()=110 find the other values.
Applications of theory of Attributes:-
Consistency of data:-
Rule:-
- No class frequency can be negative. Frequency of every class (OR)
- No class frequency can be greater than N. (Each class frequency )
Q,, Determine consistency in the given data: (AB) = 70, (B) = 40, () = 60, B = 100.
| Attributes | A | Total | |
|---|---|---|---|
| B | (AB) 60 | (B) 40 | (B) 100 |
| (A) 70 | () 60 | () 130 | |
| Total | (A) 130 | () 100 | N 230 |
Conclusion: No frequency is negative, nor any frequency is greater than N. Data is consistent
Q,, Find out if the data is consistent or not. values given are N = 300, (A) = 200, B = 180, (AB) = 170.
| Attributes | A | Total | |
|---|---|---|---|
| B | (AB) 170 | (B) 10 | (B) 180 |
| (A) 30 | () 90 | () 120 | |
| Total | (A) 200 | () 100 | N 300 |
Data is consistent
Q:- N = 400, (A) = 300, (B) = 280, (AB) = 170. Determine consistency. Sol:-
| Attributes | A | Total | |
|---|---|---|---|
| B | (AB) 170 | (B) 110 | (B) 280 |
| (A) 130 | () -10 | () 120 | |
| Total | (A) 300 | () 100 | N 400 |
Data Inconsistent
H/W Q,, (A) = 200, (B) = 100, (B) = 80, N = 300. Determine consistency.
Trick :- To check consistency of data. Check for the ultimate frequencies, If any ultimate frequency is "negative", the data is inconsistent otherwise not.
Q,, If (AB)=100, (B)=40, (A)=30 & ()=60. Find if Data is consistent or not. Sol:- No ultimate frequency is negative. Data is consistent.
H/W:- Q:- Determine the consistency of Data. (A)=220, (B)=130, ()=110, (A)=180. Find consistency of Data.
2. Independence And Association:-
- Attributes are said to be Independent if there does not exist any relation between them. e.g:- Gender and success, Beauty and Intelligence.
Two attributes are said to be associated if they are related in one way or other. e.g:-
- Positive Association: Present or Absent together. e.g.: unemployment & poverty.
- Negative Association: one is present & another absent. e.g.: Education & Ignorance.
Methodology to check Association and Independence of Attributes:-
(1) Proportion Method :-
(i) Independent (ii) Positive Association (iii) \frac{(AB)}{(B)} < \frac{(A\beta)}{(\beta)} Negative Association
Note: (,), (A,), (,B) are also Independent.
Q,, (AB) = 100, (B) = 10, (A) = 150 & () = 15. Find out how A & B are Associated? Sol:- Independent.
Q,, (AB) = 200, (B)= 10, (A) = 250 and () = 25 Sol:- Positively Associated.
H/W Q// (AB) = 210, (B) = 10, (A) = 310, () = 10. Find association?
(2) Comparison Method :-
(i) Independent (ii) Positively Associated (iii) (AB) < \frac{(A) \times (B)}{N} \rightarrow Negatively Associated
Q,, Type of Association in the given data N = 106, (A) = 70, (B) = 36, (AB) = 20. Sol:-
\therefore (AB) < \frac{(A) \times (B)}{N} \rightarrow Negatively associated.
3:- Yule's coefficient of Association method:-
Representation of values of Q. (i) If Q lies between 0 to 1 Positively Associated (ii) Q lies between -1 to 0 Negatively Associated (iii) Q = 0 Independent (iv) Q = 1 completely (Perfectly) Associated (v) Q = -1 completely (Perfectly) Disassociated.
Completely
Disassociated Negatively Associated Positively Associated Completely
|----------------------|-----------------------------|----------------------| Associated
-1 -0.5 0 0.5 1
^ ^ ^ ^ ^
strong weak Independent weak strong
(AB) or (αβ) (Aβ) or (αB)
= 0 = 0
Q:- If (AB) = 300, (B) = 15, (A) = 350 and = 25. Find the Association between A & B. Sol:- Since the values given are (AB), (B), (A) & (), we can use proportion method. Positive Association b/w A & B.
Q,, If N = 100, (A) = 50, (B) = 30 & (AB) = 40. Find Association between A & B. Sol:- Since the values given are (AB), (A), (B) and N Comparison method will be time saving. Here Positive Association between A and B.
Question Asked in PAA (JKSSB) - Year 2020
"N = 200, attributes A = 100 and B = 140 are INDEPENDENT. Find the ultimate class frequency (AB)." a) 60 · b) 70 ✅ · c) 80 · d) 90 ⭐ Working: for independent attributes (AB) = (A) × (B) ÷ N = (100 × 140) ÷ 200 = 70 ⭐ This exact formula has now been asked in BOTH 2020 and 2022. Learn it cold — it is the single highest-frequency formula of this topic.
Q:: If N = 200, which of the options match the ultimate class frequencies, given that there are two independent attributes A = 100, B = 140? (A) 60 (B) 70 (C) 80 (D) 90 Sol:- Given values are N, (A), (B) It is also given in question that Attributes A & B are independent. As per the given values, comparison method is to be used. ? Condition for independent Attributes in Comparison Method.
3:- Yule's coefficient of Association method:-
(Repeated Page/Content)
Representation of values of Q. (i) If Q lies between 0 to 1 Positively Associated (ii) Q lies between -1 to 0 Negatively Associated (iii) Q = 0 Independent (iv) Q = 1 completely (Perfectly) Associated (v) Q = -1 completely (Perfectly) Disassociated.
Completely
Disassociated Negatively Associated Positively Associated Completely
|----------------------|-----------------------------|----------------------| Associated
-1 -0.5 0 0.5 1
^ ^ ^ ^ ^
strong weak Independent weak strong
(AB) or (αβ) (Aβ) or (αB)
= 0 = 0
Q:- (AB) = 60, () = 20, A = 80, (B) = 30. Find the Association. Sol:- Since the given values are (AB), (), (A) and (B) Yule's coefficient method will be used.
Q2:- (AB) = 70, () = 30, (A) = 0, (B) = 60. Find Association. Sol:-
Q:- Find the Association between Literate Husband and Literate wife. Literate husband with Literate wife = 80 Literate husband with illiterate wife = 30 Illiterate husband with Literate wife = 100 Illiterate husband with illiterate wife = 60.
Sol:- A : Literate Husband B : Literate Wife : Illiterate Husband : Illiterate wife
(AB) = 80, (A) = 30, (B) = 100, () = 60
H/W Q!- (AB) = 90, () = 40, B = 60 and A = 35. Find association?
🔷 TOPIC (viii) — THEORY OF INDEX NUMBERS
Syllabus topic (viii) of 10 · runs until Topic (ix)
Covers: Characteristics · problems in construction · types · simple (unweighted) index · quantity index · value index · weighted aggregative — Laspeyres, Paasche, ⭐ Fisher's Ideal · weighted average of price relatives · CPI · WPI · CPI vs WPI · ⭐ Tests of Adequacy (Unit, Time Reversal, Factor Reversal, Circular) 🎯 Asked in the papers: 2024 · Q78 — Fisher Ideal Value Index — the hardest question on the 2024 paper
Index Numbers
Statistical measure that shows changes in variables with respect to time, geography or other characteristics.
- First time calculated by Italian statistician "Giovanni Rinaldo carli" Calculated ratio of prices for (grain, wine and oil). 1500 and 1750
- Also known as "Economic Barometer"
Characteristics:-
(1) Expressed in percentage form. (2) Relative or comparative measurement of a group of commodities. (3) Represent the specialised averages. (4) e.g; consumer Price Index, cost of living index, Industrial production Index.
Problems in construction of Index Numbers:-
- Purpose of Index number should be pre-defined, every index number has its specific use.
- Selection of Base year:-
- Base year should not be too near or too far.
- Should be calamity free.
- Selection of commodities.
- Choosing the source of data.
- Choice of average.
- Choice of calculation method.
Limitations of Index Numbers:-
- Index numbers give approximate values.
- Based on samples only.
- Quality of the products is not taken in consideration.
- For every purpose, different index numbers are to be constructed.
Types of Index Numbers:-
- Price Index numbers
- Quantity Index Numbers.
- Value Index numbers
- Price Index measures change in price b/w Base year and current year. Whole sale Price Index Number Retail Price Index Number
- Quantity Index numbers show average changes in quantities, produced, consumed or sold. e.g Imports, Exports, production in Industries etc.
- Value Index Number represent the product of commodity and the quantity.
Methods of constructing Index Numbers
graph TD
A[Index Numbers] --> B["==Unweighted / Simple Index Numbers=="]
A --> C["==Weighted Index Numbers=="]
B --> D["==Simple Aggregative Method=="]
B --> E["==Simple Average of Price Relatives=="]
C --> F["==Weighted Aggregative Method=="]
C --> G["==Weighted Average of Price Relatives=="]
Simple Index Numbers:-
Each item has got the same weight therefore no individual weights are assigned. (i) Simple Aggregative method (ii) Simple Average of Price Relatives.
(i) Simple Aggregative Method. Procedure:- (i) Add all the current year prices of various commodities. (ii) Add all the base year prices of various commodities.
Formula:- Price of "1" on "0"
- Index number of current year
- total (sum) of the current year prices.
- total (sum) of the Base year prices.
Q,,
| Commodity | Base year Price (Rs.) | Current year Price (Rs.) |
|---|---|---|
| A | 10 | 20 |
| B | 20 | 30 |
| C | 25 | 35 |
| D | 30 | 40 |
| E | 35 | 45 |
Sol:- Price index of current year is 141.66 (or) Price of current year has increased by 41.66%.
Quantity Index
Formula:-
- quantity index no. of current year.
- sum total of current year quantities.
- sum total of Base year quantities.
Q,, Calculate the quantity index number for 2019 when the base year is 2011.
| Commodities | Quantity (2019) (tons) | Quantity (2011) (tons) |
|---|---|---|
| A | 10 | 5 |
| B | 20 | 10 |
| C | 30 | 15 |
| D | 20 | 10 |
| E | 15 | 5 |
Sol:- Quantity index for current year (2019) = 211.11 (or) Quantity of current year has increased by 111.11.
(ii) Simple Average of Price Relatives Method
Procedure:- (i) calculate relative price of current year (current year price by Base year price) (ii) obtain the sum of Relative prices / Price Relative (iii) Divide the sum of Relative prices by total number of commodities.
Q:- Using Price Relative Method, for the year 2020 find out index values from the given data.
Sol:-
| Name | 2010 Price (Rs.) | 2020 Price (Rs.) | Price Relative = |
|---|---|---|---|
| A | 10 | 20 | 200 |
| B | 15 | 25 | 166.67 |
| C | 20 | 30 | 150 |
| D | 25 | 35 | 140 |
| E | 20 | 30 | 150 |
N = 5 | 806.67
The Price index for the year 2020 is 161.33 or there is increase in the price of 2020 by 61.33% compared to 2010.
H/W:- calculate the Quantity Index Number for the Previous question Assuming prices in (Rs.) as quantity in quintals.
Formula:-
The Price index for the year 2020 is 161.33 or there is increase in the price of 2020 by 61.33% compared to 2010.
H/W:- calculate the Quantity Index Number for the Previous question Assuming prices in (Rs.) as quantity in quintals.
Formula:-
Ans:- Same as above question.
Value Index Number :
(or) Prices increase by 8.33% from base year.
| Commodity | Base year | Current year | $P_1 Q_1$ | $P_0 Q_0$ | ||
|---|---|---|---|---|---|---|
| Quantity $q_0$ | Price $P_0$ | Quantity $q_1$ | Price $P_1$ | |||
| A | 2 | 3 | 2 | 4 | 8 | 6 |
| B | 3 | 1 | 3 | 2 | 6 | 3 |
| C | 4 | 2 | 1 | 3 | 3 | 8 |
| D | 5 | 3 | 2 | 3 | 6 | 15 |
| E | 2 | 2 | 4 | 4 | 16 | 4 |
| $\sum P_1 q_1 = 39$ | $\sum P_0 q_0 = 36$ | |||||
Weighted Index Numbers :-
Appropriate weights are assigned to various commodities to show their relative importance.
1. Weighted Aggregative Method :-
There are lot of methods to calculate weighted index numbers viz; Laspeyre's, Paasche's, Fisher's, Kelley's, Bowley's, Dorbish and few more methods.
Laspeyre's Method :-
It is used to compare the expenditure of basket of commodities of Base year and current year.
Price Index * Quantities of base year are taken as weights.
Quantity Index * Prices of Base year as weights.
Paasche's Method :-
Used to know the cost of basket of commodities in current year when the same basket cost ₹100 in Base year.
Price Index
Quantity Index
Quantities of current year and Prices of current year are used as weights.
Fisher's Method :-
"Commodity Y: price index = 10 for 1993 with base 1990; quantity index = 0.5 for 1990 with base 1993. Find the Fisher Ideal VALUE index for 1993 with base 1990." a) 5 · b) 10.5 · c) 20 ✅ · d) 15 ⭐ Step 1 — spot the reversal. The quantity index is given the wrong way round (1990 on base 1993). Flip it with the Time Reversal Test, : ⭐ Step 2 — apply the Factor Reversal Test, Value Index = Price Index × Quantity Index: ⚠️ This is why the two tests below matter. The question is unsolvable unless you know that Fisher satisfies BOTH — the whole item is really a test of the Time Reversal and Factor Reversal rules, not of arithmetic.
Also called as Ideal method.
Price Index:-
- It is the geometric mean of Laspeyre's and Paasche's method.
- Both Base and current year Quantities are used as weights.
Quantity Index :-
- Both Base year and current year Prices are used as weights.
| Laspeyre's (Base) | Paasche's (Current) | Fisher's method | |
|---|---|---|---|
| Price | $P_{01} = \frac{\sum p_1 q_0}{\sum p_0 q_0} \times 100$ | $P_{01} = \frac{\sum p_1 q_1}{\sum p_0 q_1} \times 100$ | $P_{01} = \sqrt{\frac{\sum p_1 q_0}{\sum p_0 q_0} \times \frac{\sum p_1 q_1}{\sum p_0 q_1}} \times 100$ |
| Quantity | $q_{01} = \frac{\sum q_1 p_0}{\sum q_0 p_0} \times 100$ | $q_{01} = \frac{\sum q_1 p_1}{\sum q_0 p_1} \times 100$ | $q_{01} = \sqrt{\frac{\sum q_1 p_0}{\sum q_0 p_0} \times \frac{\sum q_1 p_1}{\sum q_0 p_1}} \times 100$ |
Q,, Calculate Price and quantity Index numbers using Paasche's, Lapeyere's and Fisher's method.
| Commodities | $p_0$ | $q_0$ | $p_0 q_0$ | $p_1$ | $q_1$ | $p_1 q_1$ | $p_1 q_0$ | $p_0 q_1$ |
|---|---|---|---|---|---|---|---|---|
| A | 2 | 3 | 6 | 6 | 2 | 12 | 18 | 4 |
| B | 3 | 4 | 12 | 3 | 4 | 12 | 12 | 12 |
| C | 4 | 2 | 8 | 2 | 3 | 6 | 4 | 12 |
| D | 5 | 3 | 15 | 4 | 7 | 28 | 12 | 35 |
| $\sum = 41$ | $\sum = 58$ | $\sum = 46$ | $\sum = 63$ |
Laspeyre's method:
Price Index |
Quantity Index |
Paasche's Method :-
Price Index :- / Fallen
Quantity Index (Correction: formula written in notes uses for weights) /
Fisher's Method :
Price Index or
Quantity Index or
(2) Weighted Average of Price Relatives Method:
Procedure (i) Calculate Price Relatives of current year
(ii) calculate the value weights
Price Index
Q,, calculate the weighted Average Price Relatives Index for given data.
| Commodities | Price 2011 | Quantity | Price 2018 | Price Relative | ||
|---|---|---|---|---|---|---|
| A | 8 | 20 | 10 | 160 | 20,000 | |
| B | 6 | 10 | 8 | 60 | 7999.8 | |
| C | 4 | 8 | 6 | 32 | 4800 | |
| D | 2 | 6 | 4 | 12 | 2400 | |
(or) Increase in prices of 2018 on 2011.
Consumer Price Index (CPI) :- (weighted)
Also known as:-
- Real Price Index Number.
- Cost of Living Index number.
- Price of Living Index Number.
- Retail Price Index Number.
CPI is used to measure the price of basket of goods and services of a particular class or region at particular point of time in comparison to Base year.
Applications of CPI:
(1) To know increase or decrease in cost of living. (2) Used by government to make salary and (D.A). (3) Used to find purchasing power of money and real income/wages.
(4) Used by government in framing price policy and income policy.
Types of CPI:
- (CPI - IW) - Industrial workers (1982 B.Y) 2001 2016 [Ministry of Labour, Labour Bureau (Shimla)]
- (CPI - AL) - Agricultural Labours (1986-87 B.Y)
- (CPI - RL) - Rural Labours (1986-87 B.Y)
- (CPI - UNME) - Urban Non-Manual employees (1984-85 B.Y) (CSO) Now (NSO) [MOSPI - Ministry of Statistics & Program Implementation]
2011 - CPI (R), CPI (U), CPI (Combined)
The CPI (Rural / Urban / Combined) series above ran on the 2012 base year. That is no longer current.
⭐ The new CPI series has base year 2024, first released on 12 February 2026
- Built on the Household Consumption Expenditure Survey (HCES) data
- Published by MoSPI / NSO
- ⭐ Remember the pair for 2026: CPI base = 2024 · WPI base = 2022-23
(The sub-index base years above — CPI-IW 2016, CPI-AL/RL 1986-87 — are separate series and remain as stated.)
Problems in construction of CPI:-
- Difference in price of goods and services.
- Difference in the standard of living.
- Choice of Base year.
- Difference in proportion of expenditure.
Methods of construction of CPI:-
- Aggregate Expenditure Method
- Family Budget Method
(1) Aggregate Expenditure method / Weighted Aggregate Method :- Quantities of base year are taken as weights ** Laspeyere's method
(2) Family Budget Method / Weighted Average of Price Relatives :- Expenditure in the base period are taken as weights. Weights "" =
Q,, Find the cost of living index in the given Data.
| Commodities | |||||
|---|---|---|---|---|---|
| A | 5 | 4 | 10 | 40 | 20 |
| B | 3 | 2 | 6 | 12 | 6 |
| C | 2 | 3 | 4 | 12 | 6 |
| D | 1 | 4 | 2 | 8 | 4 |
Sol:- Cost of Living =
** Based on Lapeyre's Price Index.
Q,, Data about the middle class family is as follows.. Expense on food 30%, Rent 15%, clothing 20%, Fuel 10%, others 25% on base year. Price (₹) 2001: 100, 20, 70, 20, 40 Price (₹) 2011: 90, 20, 140, 15, 60 Find Cost of living.
Sol:-
| Commodities | Expenses (%) (W) | WP | |||
|---|---|---|---|---|---|
| Food | 30 | 100 | 90 | 2700 | |
| Rent | 15 | 20 | 20 | 1500 | |
| Clothing | 20 | 70 | 140 | 4000 | |
| Fuel | 10 | 20 | 15 | 750 | |
| Others | 25 | 40 | 60 | 3750 | |
Whole sale Price Index:-
It measures the general changes in the whole sale price of goods in the country.
- Based on the commodities produced and distributed (first stage of transaction).
- Rise or fall in prices at wholesale level spills over to the retail level after lag.
- Published by Economic Advisor, Ministry of Commerce and Industry.
- First time published 10th January 1942 (1939 B.Y).
- 7th revision is with Base year (2011-2012) chaired by Dr. Sumitra Chaudhari. Education, Health, etc not included. * taxes are also not included.
- WPI Food Index (CSO) separately presented.
- 697 commodities included in WPI.
The 2011-12 base year and the 697 commodities above are the OLD series. They were superseded before your exam.
| Old series | ⭐ NEW series | |
|---|---|---|
| Base year | 2011-12 | ⭐ 2022-23 |
| Effective from | — | ⭐ June 2026 (with the May 2026 indices) |
| Number of items | 697 | ⭐ 957 |
| Weights basis | — | GVO-based; renewable energy now included |
⭐ A new PRODUCER PRICE INDEX (PPI) was introduced alongside it, intended to eventually replace the WPI. Released by DPIIT, Ministry of Commerce & Industry (unchanged).
⚠️ The three-group weights were also rebased, so the 22.60 / 13.20 / 64.20 split below is superseded. Reported figures for the new series put Manufactured ≈ 65%, Primary Articles ≈ 20%, Fuel & Power ≈ 5% — but confirm the exact percentages against the DPIIT release before memorising them. The base-year change is the part that gets asked.
WPI Basket
graph TD
A[WPI Basket] --> B["Primary Articles<br>(22.60%)<br>117 items<br>Rice, wheat, fruits etc."]
A --> C["Fuel & Power Articles<br>(13.20%)<br>16 items<br>Power, coal petroleum etc."]
A --> D["Manufactured Articles<br>(64.20%)<br>564 items<br>oil (edible), sugar, chemicals etc."]
Uses of WPI:-
(1) Estimation of Inflation:
(2) Estimation of Monetary value and Real value. (3) Used for estimating GDP by CSO. (4) Used by Business contractors (Demand & supply). (5) By Global investors for investment decisions.
Method for Calculation:
Stage 1: Elementary price Indices using Jevon's index (Geometric mean for price). Stage 2: Elementary aggregated using Laspeyre's index formula.
WPI (Vs) CPI
| WPI | CPI |
|---|---|
| 1. Released by Office of Economic Advisor (Ministry of Commerce & Industry) | 1. NSO (Ministry of statistics and Program Implementation) |
| 2. Measures Goods only | 2. Both Goods & services. |
| 3. Items :- 697 | 3. Items — 448 (Rural Basket) 460 (Urban Basket) |
| 4. Base year :- 2011-2012 | 4. Base year: 2012. |
5. 3 categories
| 5. Many categories
|
Tests of Adequacy Index Numbers
(1) Unit Test (2) Time Reversal Test (3) Factor Reversal Test (4) Circular Test (extention of TRT)
(1) Unit tests Index number formulae should be independent of the units in which prices or quantities are used. * Satisfied by all index methods except simple (unweighted) aggregative method.
(2) Time Reversal Test:- Interchanging of time subscripts of price/quantity gives the reciprocal of the original formula. (Base year '1', Base year '0') * Not satisfied by Laspeyre & Pasche.
(3) Factor Reversal Test:- If and factors in price/quantity index formula are interchanged so that a quantity/price index formula is obtained, the product of two indices should give true value ratio. * Satisfied by Fisher Index only.
(4) Circular Test:- (extension of TRT) If the index for year 2018 is based on 2017 and Another index for 2017 based on 2016 then index for 2018 with base year 2016 should be directly obtained. * Satisfied by:
- simple geometric mean of price relatives
- Kelly's fixed Base method (simple)
🔷 TOPIC (ix) — DEMOGRAPHY — CENSUS, ITS FEATURES AND FUNCTIONS
Syllabus topic (ix) of 10 · runs until Topic (x)
Covers: Demography & demographic transition · the population cycle · types of demography · four stages of Indian demographic history · Census — features & functions · background of the Census of India · Census 2011 data · ➕ added: national headline figures and Nagaland's negative growth 🎯 Asked in the papers: 2024 · Q79 — a feature of the census → Simultaneity · 2022 · Q40 — highest negative decadal growth → Nagaland · 2022 · Q36 — demographic sex ratio
Demography - census, its features and functions
Demography: Derived from two Greek words
- "Demos" — "the people"
- "Graphy" — Recording / writing / measuring something.
Study of population of a country or place. Term "Demography" Guillard (1855) Father of "Demography" John Graunt (Natural & Political observation (1662))
Deals with Fertility, Marriage, Mortality, Migration, Social mobility. Mainly deals with: Migration, Birth rate, Death rate.
Demographic Transition :-
Shift from high birth rate and high infant death rate to low birth rate and low infant death rate with minimum education & technology.
Demographic Transition Model (Population cycle) :-
It shows the growth rates in populations and effect on population. Divided in four stages:
Stage 1 - High Fluctuating (High stationary) Both Birth rate and Death rate are both high. * Population growth slow and fluctuating.
| BR Reasons | DR Reasons |
|---|---|
| • Family planning • Religious beliefs • Child - Economic Asset | • Diseases • Famine • clean water & sanitation • war • Education |
| e.g.: Britain in 18th Century & (LEDC's today) |
Stage 2 :- Early Expanding : Birth rate remains high, Death rate falls. * Population begins to rise steadily.
| Death Rate Reasons! - |
|---|
| • Improved H.care • Improved Hygiene • Improved Food Production • Less infant mortality rate. |
| e.g.: Britain in 19th Century, Bangladesh, Nigeria. |
Stage 3 :- Late Expanding :- Birth rate starts to fall. Death rate starts to fall. * Population rising.
| Reasons:- |
|---|
| • Family planning • Low infant mortality. • Increased standard of living • changing status of women. |
| e.g.: Britain in Late 19th & early 20th century ; China, Brazil, India. |
Stage 4 :- Low Fluctuating :- (low stationary) Birth rate and Death rate both low. * Population steady. e.g.: USA, Sweden, Japan, Britain.
Stage 5 :- Declining stage. BR < DR (Germany, Hungary).
- Model Assumes all countries pass through all four stages.
- Assumes fall in death rate in stage (2) is because of Industrialisation.
- Countries that grew as a consequence of emigration did not pass through early stages of model (USA, Canada, Australia).
Demographic Dividend: When the working population of a country is more than non-working population. A.K.A (Demographic Bonus) Demographic Burden: Working population is less than non-working population. Doubling time: It is number of years required to double the population of a country / Area at a given growth rate.
Father of "Demographic studies" — Karl Marx
Types of Demography :-
1. Formal Demography :- Mathematical study / statistical Analysis - of numbers (population). e.g.:- No. of males / No. of Females, No. of employed / unemployed people.
2. Social Demography :- Deals with the changes and consequences due to population. e.g.: Birth rate, Death rate, emigration etc.
Four distinct stages of Indian Demography History :-
- Period of stagnant Population (1901-1921) (Population more or less stagnant)
- Period of steady Growth (1921-1951) (More than 1% Growth rate / year) Fertility Mortality.
- Period of Rapid High Growth (1951-1981) (Growth rate > 2% / year) Period of population Explosion
- Period of High Growth with definite signs of slowing down. (1981-2011) 2.22% (1971) 1.64% (2011) world (1.23% GR)
* 1921 Year of Great Demographic Divide. (After this, there was continuous increase in population).
| Antinatalist policy | Vs | Pronatalist Policy |
|---|---|---|
| Govt policy to slow down population | Govt Policy to Increase the population of country. | |
| e.g; China (one child policy) | e.g; Hungary, Japan etc. |
Census
Complete enumeration method which systematically records information about the members of given population. e.g; Housing, culture, Business, Agriculture etc.
Features of Census :-
"One of the features of the census is:" a) Defining boundaries · b) Simultaneity ✅ · c) Employment of census-workers · d) Homogeneity ⭐ Answer: B — Simultaneity. Straight recall from the list below. ⚠️ Note how close the distractors are: "defining boundaries" sounds like Defined Territory and "homogeneity" sounds like Universality. Learn the exact words, not the idea.
(1) Sponsorship :- National Govt / State Govt / Local Bodies. (2) Defined Territory :- Clear Demarcation of Area. (3) Well Defined Periodicity :- Done after regular intervals to compare and Analyse Information. (4) Universality :- Include each person without the territory without redundancy. (5) Compilation and Publication :- Filtering and presentation of useful data and publishing it for potential users. (6) Dissemination :- Getting right data for right people. (7) International simultanity :- compared with other countries.
[!fix] ⚠️ DEFINITION CHECK — "Simultaneity" 2024·Q79 asked for a feature of the census and the answer was "Simultaneity." The gloss above ("compared with other countries") is not the standard meaning. ⭐ Simultaneity = every person is enumerated with reference to the SAME point in time (a single reference moment), so the count is a true snapshot. That is what makes it a census feature. ❗ Also absent from this list: "Individual Enumeration" — each person is recorded separately — which is one of the six standard UN census features and a likely option.
Functions :-
(1) For Administrative purpose and Policy Making: Used to analyse population of area for employment programs, housing, Education, social welfare schemes, economic aspects etc. (2) For Research purpose: Used the solution and interpretation of scientific problems. (3) For social and economic planning of the country: * Gender classification, Urban Vs Rular classification etc. (4) Utility in Business and Industrial sector: Demand and Supply Analysis.
Background of Census of India:
(U.S.A (1790) First modern census)
- In 1830 (Dacca) First census was done (Henry Walter) Father of undivided India
- 1872 First Indegenious census (Lord Mayo) V.G of India
- 1881 First synchronous population census. (Lord Rippon) V.G of India 1st. Census commissioner W.C Plowden (Father of Indian census 1931 First caste based census of India. (Hutton - commissioner) Caste in India (1946).
1961 MHA RGCCI (Registrar General & census commissioner of India) Present - Vivek Joshi
15th census (1872) (7th after independence) Census 2011 (Our census, our future) RGCCI - Dr. Chandramauli Mascot: Women Enumerator (stick figure)
- No. of states and UTs (35)
- Districts 640 (Increased by 47 from 2001)
- Towns 7,933 ( by 2772 from 2001)
- No. of villages 6,40,930 ( 2342 from 2001)
- Total Population 1,21,05,69,573 Urban (68.8%) Rural (31.2%)
- Child Sex Ratio (0-6 year) 919 Rural (923) Urban (905)
- Sex Ratio 940 / thousand males.
- Density (Pop) 382 person / .
- Decadal pop. Growth 17.64%.
- India's total pop. of world 17.5%.
- Literacy rate 74.04%.
- India pop. USA + Indonesia + Brazil + Pakistan + Bangladesh.
Population wise Rank:-
China (19.4% WP) India (17.5% WP) USA (4.5%) Indonesia Brazil
The line above shows India overtaking China in 2030. That is wrong.
⭐ India overtook China in APRIL 2023 and is now the world's most populous country
- UN estimate at the crossover: India 1,425,775,850 people
- China peaked at 1.426 billion in 2022 and has been falling since
- ⭐ Correct order today: INDIA > China > USA > Indonesia > Pakistan
(India's 7th rank by AREA, given below, is still correct.)
Area wise Rank:-
Russia Canada China US Brazil Australia India (7th rank)
The notes stop at Census 2011. The next census has been announced and its first phase is happening as you revise.
| Phase I — Houselisting & Housing Census | ⭐ April to September 2026 (in progress) |
| Phase II — Population Enumeration | ⭐ March 2027 |
| Reference date | 1 March 2027 (1 October 2026 for snow-bound areas — J&K, Ladakh, Himachal, Uttarakhand) |
| Significance | ⭐ India's FIRST DIGITAL census · first since 2011 · will include caste enumeration |
| Which census | The 16th census of India, 8th since Independence |
⭐ The J&K angle: the 1 October 2026 reference date for snow-bound areas covers J&K — a very likely J&K-GK question. ⚠️ Census 2011 remains the latest COMPLETED census, so every 2011 figure below is still the correct exam answer.
➕ Census 2011 — the national headline figures
These were NOT in the original notes, which jump straight to the state-wise ranks. The paper asks the headline numbers too.
| Figure | Value |
|---|---|
| Total population | 121.09 crore (1.21 billion) |
| ⭐ Decadal growth (2001–11) | 17.70% |
| Population density | 382 per km² |
| ⭐ Sex ratio | 943 females per 1,000 males (the 940 above is the provisional figure) |
| Child sex ratio (0–6) | 919 |
| Literacy rate | 74.04% (M 82.14 · F 65.46) |
| Census year / which census | 2011 — the 15th national census, 7th after Independence |
Other decadal-growth extremes (2001–11):
- Highest growth — state: Meghalaya (27.9%) · UT: Dadra & Nagar Haveli (55.9%)
- Lowest / negative — state: ⭐ Nagaland (−0.58%)
⬇️ Your original notes resume below. ⬇️
As per 2011 census!
Population wise
Biggest-state :- UP (16.49% of Total Population) Maharashtra Bihar Smallest-state :- Sikkim (0.05%) Biggest-UT :- Delhi Smallest-UT :- Lakshadeep
Mnemonic:
U S Se Dil Lagi bhool jani Padgi
Area wise:
Largest-state :- Rajasthan (3,42,240 ) Smallest-state :- Goa (3702 ) Largest-UT :- Andaman & Nicobar (8249 ) Smallest-UT :- Lakshadeep (32 )
Mnemonic:
Roz Ghar Ana Ladke
Sex Ratio (940) overall:
Highest-state :- Kerala (1084) Tamil Nadu (995) Lowest-state :- Haryana (877) Highest-UT :- Puducherry (1038) Lowest-UT :- Daman & Diu (618)
Mnemonic:
Kash Har Ghar mein Daughter hoti
[!fix] ⚠️ FIGURE CHECK — sex ratio The heading above gives the overall 2011 sex ratio as 940. That was the PROVISIONAL figure; the FINAL Census 2011 figure is 943 females per 1,000 males. Use 943 in the exam. ⭐ Also note the convention trap: 2022·Q36 gave the ratio the other way round — (men ÷ women) × 100 = 96.70. Read the options to see which convention is intended. ❗ Missing here: Census 2011's standout fact — Nagaland was the ONLY state with NEGATIVE decadal growth (−0.58%). 2022·Q40 asked exactly this.
Literacy Rate (74.04%):
M 82.14% F 65.46%
Highest-state :- Kerala (93.91%) Mizoram (91.58%) Smallest-state :- Bihar (63.82%) Highest-UT :- Lakshadeep (92.28%) Lowest-UT :- Dadra & Nagar Haveli (77.65)
Mnemonic:
Koun Banega Lakhpati - Haveli
Urbanization wise:
Most-state :- Goa (62.17%) Least-state :- Himachal Pradesh (10.04%) Most-UT :- Delhi (97.50%) Least-UT :- Andaman & Nicobar (35.67%)
Mnemonic:
Ghar Ho tou Delhi k Andhar
District wise:-
Population wise:
- Biggest - Thane (Mumbai)
- Lowest - Dibang valley
Sex Ratio:-
- Highest :- Mahe (1184) Puducherry
- Lowest :- Daman (534) (D&D)
Literacy Rate:
- Highest :- Serchipp (Mizoram) 98.76%
- Lowest :- Ali Rajpur (M.P) (36.10%)
Census day - 1st April (FY) National Census Day - 9th Feb Population Day - 11th July
🔷 TOPIC (x) — VITAL STATISTICS
Syllabus topic (x) of 10 · runs until the end of the syllabus
Covers: Objectives & sources · key terms (density, CBR, CDR, natural increase, life expectancy, IMR, neonatal, MMR) · Fertility — CBR, GFR, SFR, TFR · Mortality — CDR, SDR, STDR · ⭐ GRR · ⭐ NRR · ➕ added: NRR ≤ GRR, ASFR modal age, MMR denominator conventions 🎯 Asked in the papers: 2024 · Q80 — match the measures of mortality/fertility
Vital Statistics
Study and measurement of vital factors related to health and growth of community. It includes:-
- Marriage
- Births
- Deaths
- Sickness
- Migration and so on.
Objectives:-
- To implement and evaluate health related schemes (National Health Programs).
- To determine community health related infections, epidemics and find solution.
- To use the data as primary tool in research activities.
- Administrative purposes.
"Branch of Biometry that deals with data and law of human mortality, morbidity & demography".
Sources of Vital Statistics
- Civil Registration system.
- National Surveys.
- Health Surveys.
- Sample Registration system.
Key Terms
- Population Density: Population per square Area / Region (Pop. / )
- Crude Birth Rate: Number of Births / 1000 people.
- Crude Death Rate: Number of Deaths / 1000 people.
- Natural Increase: (CBR - CDR) per year in percentage (%).
- Life Expectancy: How long people in certain countries are expected to live.
- Replacement Rate: The rate that is needed to replace both parents.
- Age dependency: Percentage of population that depend on economic support (e.g; pension for elderly, school for young).
- Age Dependency Ratio:
- Abortion Ratio: Number of Abortions per hundred pregnancies.
- Abortion Rate: Number of Abortions per thousand women (15-44) age.
-
Emigration Rate: Number of Emigrants per thousand population per year.
-
Immigration Rate: No. of Immigrants per thousand population per year.
-
Brain Drain: Emigration of highly trained, skilled people from country.
-
Effective literacy rate: Number of Literates with age more than or equal to 7 to the population with age more than or equal to 7 in year.
-
Infant mortality Rate:
-
Neonatal Mortality Rate:
-
Post-neonates Mortality Rate:
-
Maternal Mortality Rate:
Fertility
"Natural capacity to give birth to an offspring"
Fecundity:- capacity to bear (how many) child births.
Fertility Rate / Natality Rate / Birth Rate:-
Measure of rate of growth in population due to births in a specific period.
- Expressed in per thousand women per year.
- Child Bearing Age (India) (15 - 49) years.
Methods to calculate Birth rate:-
- Crude Birth Rate (CBR)
- General Fertility Rate (GFR)
- Specific Fertility Rate (SFR)
- Total Fertility Rate (TFR)
(i) Crude Birth Rate (CBR)
Q,, No. of live births in particular area is 3000, If the total population is 3 lac. Calculate crude Birth rate. Sol:-
Q,, Find the crude birth rate for given data.
| Age (yrs) | No. of Females | No. of Males | No. of Births | ASFR |
|---|---|---|---|---|
| 0 - 14 | 1000 | 1500 | - | |
| 15 - 25 | 2000 | 2000 | 50 | |
| 26 - 35 | 1000 | 1000 | 100 | |
| 35 - 40 | 3000 | 1500 | 50 | |
| 41 - 49 | 1500 | 3000 | 50 | |
| 50 - Above | 1000 | 1000 | - | |
Sol:-
(ii) General Fertility Rate (GFR)
For the Above Data:
(iii) Specific Fertility Rate (SFR):-
Fertility Rate taking the factors which affect it like; Marriage, Age, migration etc.
ASFR =
(iv) Total Fertility Rate (TFR)
(a) For equal intervals = (b) For unequal intervals =
| Age(yrs) | No. of Females | No. of Males | No. of Births | ASFR | TFR = ASFR |
|---|---|---|---|---|---|
| 0 - 14 | 1000 | 1500 | - | - | |
| 15 - 25 | 2000 | 2000 | 50 | 25 | |
| 26 - 35 | 1000 | 1000 | 100 | 100 | |
| 35 - 40 | 3000 | 1500 | 50 | 16.67 | |
| 41 - 49 | 1500 | 3000 | 50 | 33.33 | |
| 50 - Above | 1000 | 1000 | - | - | |
Case (b):- unequal intervals
Quick Recap
Mortality (Death Rate):-
Number of deaths per thousand population in a particular area in specified time period.
Methods of calculating Death Rate:-
(1) Crude Death Rate (CDR) (2) Specific Death Rate (SDR) (3) Standardised Death Rate (STDR)
(1) Crude Death Rate (CDR):-
Q,, Find the Male, Female and crude Death Rate from the given data.
| Age (yrs) | Female Pop(1000) | Male Pop (1000) | No. of deaths (Total) | No. of Female Deaths | No. of male Deaths |
|---|---|---|---|---|---|
| 0 - 20 | 15 | 15 | 120 | 60 | 60 |
| 20 - 40 | 20 | 20 | 200 | 50 | 150 |
| 40 - 60 | 10 | 15 | 250 | 100 | 150 |
| 60 - 80 | 15 | 20 | 300 | 100 | 200 |
| 80 - Above | 25 | 30 | 450 | 200 | 250 |
Sol:
2. Specific Death Rate (SDR):-
3. Standardized Death Rate :- (STDR)
specific Death Rate Standard Population of age group.
Q Find the standardized Death Rate for the Given Data.
| Age (yrs) | Population | Deaths | S. population | ||
|---|---|---|---|---|---|
| 0 - 20 | 40,000 | 180 | 20,000 | ||
| 20 - 40 | 30,000 | 200 | 10,000 | ||
| 40 - 60 | 20,000 | 220 | 5,000 | ||
| 60 - Above | 10,000 | 300 | 2000 | ||
Gross Reproductive Rate (GRR):
To get better view about the rate of population, gender of new born baby is taken into consideration.
"Female children are potential future mothers and result in population increase"
Methods to calculate GRR:-
(1) If female births to "1000" are given.
Q,, Calculate the GRR for the given data:-
| Age Group | No. of female children born to '1000' women |
|---|---|
| 15 - 21 | 200 |
| 22 - 28 | 300 |
| 29 - 35 | 400 |
| 36 - 42 | 200 |
| 43 - 49 | 100 |
(2) Female Population and female births are given:- For unequal intervals
Q Calculate the GRR for the data given.
| Age | Female Pop. () | No. of Female Children () | |
|---|---|---|---|
| 15 - 20 | 1200 | 200 | |
| 21 - 27 | 2000 | 300 | |
| 28 - 35 | 2500 | 500 | |
| 36 - 41 | 3000 | 400 | |
| 42 - 49 | 1000 | 600 | |
| (per woman) |
(3) If total Fertility Rate and Male : Female ratio is given:-
Q,, If TFR = 1500 per thousand, Male : Female ratio is = 65:35 for new Born babies. Calculate GRR. Sol:- (per 1000)
Points to remember:
(i) If Population inspite of low birth rate (ii) If GRR < 1 \rightarrow Population inspite of low mortality (iii) If Population stagnant (population of females replacing themselves)
⭐ 1. NRR is ALWAYS less than or equal to GRR
Why: GRR counts the daughters a woman would bear assuming she survives her whole reproductive span — it IGNORES mortality. NRR applies actual survival rates, so some of those daughters are removed. You can only ever lose women, never gain them — so NRR can never exceed GRR. They are equal only in the impossible case of zero mortality.
| GRR | NRR | |
|---|---|---|
| Counts | Daughters per woman | Daughters per woman |
| ⭐ Mortality | ⭐ IGNORED | ⭐ ACCOUNTED FOR |
| Relation | — | ⭐ NRR ≤ GRR always |
| NRR = 1 | — | Population exactly replaces itself |
| NRR > 1 | — | Population will grow |
| NRR < 1 | — | Population will decline |
⭐ 2. ASFR has its modal value between 20 and 25
The Age-Specific Fertility Rate rises from age 15, ⭐ peaks in the 20–25 age group (the modal childbearing age), then falls away to nearly zero by 49.
⚠️ 3. The MMR denominator — know both conventions
The notes above give Maternal Mortality Rate per 1,00,000 live births — that is the official WHO / SRS convention and is correct. ⚠️ But some textbooks and question papers state MMR "per 1,000 live births" — the 2024 paper did exactly this. Read the options and take whichever convention the question offers. The idea being tested is only that MMR is measured against live births, not against total population.
a) GRR → measures the number of daughters over a lifetime b) ASFR → has modal value between 20 and 25 c) NRR → affected by mortality rates d) MMR → measured per 1000 live births Answer: a-5, b-3, c-2, d-4
⬇️ Your original notes resume below. ⬇️
Net Reproductive Rate (NRR) :
Net Reproductive Rate is the extension or modified form of GRR where mortality / survival rates is also taken in Account.
(i) If Female births, female population and survival rate is given :-
- interval
- Female births
- Survival Rate
- Female Population
Q,, Calculate the net Reproductive rate for the given Data.
| Age | Female Pop () | Female Births () | Survival rates () | ||
|---|---|---|---|---|---|
| 15 - 19 | 1000 | 100 | 0.5 | ||
| 20 - 24 | 500 | 50 | 0.6 | ||
| 25 - 29 | 1000 | 100 | 0.8 | ||
| 30 - 34 | 700 | 50 | 0.9 | ||
| 35 - 39 | 1000 | 40 | 0.7 | ||
| 40 - 44 | 500 | 90 | 0.6 | ||
| 45 - 49 | 300 | 50 | 0.5 | ||
Per Woman.
Trick: for calculating Survival Rate from Mortality Rate or Mortality Rate from Survival Rate :-
or
e.g:- If Survival Rate = 6, M.R = ?
e.g:- If M.R = 63 what is survival Rate
Important points About NRR:-
(i) If NRR > 1, Population will increase inspite of high death rate. (ii) If NRR < 1, Population will decrease inspite of high birth rate. (iii) If NRR = 1, Population will be stagnant.
Case 3 of trick!
Q,,
(Calculator = 15.73)
Q,, Number of men is 151,781,326 and number of women is 156,964,212. Which options represent demographic sex Ratio. (PAA JKSSB 2020)
Answer: b) 96.7 — as worked below. ⭐ This question has now appeared in BOTH 2020 and 2022. It is a repeat item — learn the method, not the number. ⚠️ Convention trap: this paper computed (men ÷ women) × 100. But India's official sex ratio is quoted the other way — females per 1,000 males (Census 2011 = 943). Read the options to see which convention is wanted.
(a) 97.6 (b) 96.7 (c) 98.2 (d) 95.3
Sol:-
,
(Calculator = 96.69)