Topic (i) — Primary and Secondary Data
<- Previous · Index of all topics · Next ->
Part of the full combined notes. · 2 PYQ callouts inside.
🔷 TOPIC (i) — PRIMARY AND SECONDARY DATA
Syllabus topic (i) of 10 · runs until Topic (ii)
Covers: data collection · primary data · secondary data · primary vs secondary · types of data · census vs sample and sampling error 🎯 Asked: 2024 · Q71 — secondary data is primarily sourced from?
1. Data Collection
The step after the research problem is identified and the research design is made is data collection.
Based on the approach to information gathering, data is categorised as:
graph TD
A[Data] --> B[Primary Data]
A --> C[Secondary Data]
2. Types of Data
➕ Added — not in the original notes.
2.1 By nature
| Quantitative (numerical) | Qualitative (categorical / attribute) |
|---|---|
| Measurable in numbers | Cannot be measured, only classified |
| Height, income, marks, age | Gender, literacy, blindness, religion |
| Handled by mean / median / mode | Handled by Theory of Attributes — topic (vii) |
2.2 Quantitative data splits again
| Variable | Takes | Example | Obtained by |
|---|---|---|---|
| Discrete | Only whole, separate values | Number of children, number of cars | Counting |
| Continuous | Any value in a range | Height, weight, temperature | Measuring |
2.3 By time reference
| Type | Meaning | Example |
|---|---|---|
| Time-series | One unit over many periods | India's GDP, 2010–2026 |
| Cross-sectional | Many units at one point in time | GDP of 20 countries in 2026 |
| Panel / longitudinal | Both together | GDP of 20 countries, 2010–2026 |
The quick test: count it in whole numbers → discrete · measure it to decimals → continuous · only label it → qualitative.
3. Primary Data
Data collected for the first time by the researcher himself, for a specific purpose. Example: a survey on job satisfaction of employees; health needs of a community.
| Advantages | Disadvantages |
|---|---|
| Investigator collects problem-specific data | More hectic and time consuming |
| No doubt about the quality of the data | Dealing with funding and funding agencies |
| Additional data can be obtained during the study | Ethical considerations (consent, permissions) |
| Costly / expensive | |
| Unnecessary or useless data is not included |
4. Secondary Data
a) Original research studies · b) Publicly available databases ✅ · c) Personal interviews · d) Laboratory experiments Answer: B. The three wrong options are all primary collection. Gather it yourself → primary; look it up → secondary.
Data that is already collected, produced or published by others. Example: use of hospital records or census records.
| Advantages | Disadvantages |
|---|---|
| Already collected — hassle free | Hard to get the specific data you need from it |
| Less expensive | Additional data or clarification cannot be obtained |
| Investigator is not responsible for data quality | Data may be of low quality or fabricated |
5. Primary vs Secondary Data
| # | Primary | Secondary |
|---|---|---|
| 1 | Real-time data | Past data |
| 2 | Time consuming | Quick and easy |
| 3 | Expensive | Economical |
| 4 | Available in crude form | Available in processed form |
| 5 | More accurate and reliable | Less accurate and reliable |
| 6 | Surveys, observations, experiments, questionnaires, schedules, local correspondents, interviews | Government publications, websites, books, journals, articles, internal records |
6. Census vs Sample, and Errors in Collection
➕ Added — not in the original notes.
6.1 Census (complete enumeration) vs Sample survey
| Census | Sample Survey | |
|---|---|---|
| Coverage | Every unit of the population | A part of the population |
| Cost / time | High | Low |
| Accuracy | High, if done well | Depends on the sample |
| Sampling error | NONE | Present |
| Non-sampling error | Present, often larger | Present, often smaller |
| Best when | Population is small; 100% detail needed | Population is large; units are destroyed on testing |
Testing that destroys the unit (bulbs, matchsticks, blood) must use sampling — a census would destroy the whole population.
6.2 Sampling error vs Non-sampling error
| Sampling Error | Non-Sampling Error | |
|---|---|---|
| Cause | Only a part of the population is studied | Mistakes in collection, recording, processing |
| Types | — | Measurement · coverage · non-response · response bias · processing |
| Occurs in a census? | No | Yes |
| As sample size increases | Decreases | May increase |
| Measurable? | Yes, statistically | Hard to measure |
a) It decreases as sample size increases ✅ The pair to remember: bigger sample → smaller sampling error, but possibly bigger non-sampling error.
6.3 Essentials of a good sample
Representative · Adequate (large enough) · Independent · Homogeneous · Free from bias
6.4 Sampling methods
| Random / Probability | Non-Random / Non-Probability |
|---|---|
| Simple random (lottery, random numbers) | Judgment / purposive |
| Stratified — split into homogeneous strata, sample each | Quota |
| Systematic — every k-th unit | Convenience |
| Cluster / multi-stage | Snowball |
In probability sampling every unit has a known, non-zero chance of selection — which is what makes the sampling error calculable.