Every survey, census, or research study lives or dies by the quality of its data. Ask the wrong question, measure a variable poorly, or lose track of who did not respond, and the final numbers can mislead policymakers, businesses, and researchers alike. Errors in data collection are generally split into two broad families: non-sampling errors, which can creep in regardless of whether you study a sample or the entire population, and sampling errors, which arise specifically because a survey studies only part of the universe rather than all of it. Understanding both is essential to judging how much trust to place in any dataset.
Table of Contents
- Errors of measurement: when the variable itself goes wrong
- How people themselves distort measurement
- False data and the problem of bias
- Recall lapse and problems with the sampling frame
- When the frame itself is faulty
- Non-response and processing errors
- Sampling errors: the cost of studying a part, not the whole
- Why this distinction matters
Errors of measurement: when the variable itself goes wrong
The most basic kind of non-sampling error happens at the point of measurement. If the variable being measured, say the milk yield of a cow, is not clearly defined, or if the measuring instrument and method are inconsistent, the resulting figures will be unreliable no matter how careful the rest of the survey is. One investigator might record the yield from a single milking, another might average two milkings, and a third might round off figures. None of these approaches is wrong in isolation, but using them interchangeably across a survey introduces noise that has nothing to do with the actual cows being studied.
The fix is straightforward in principle: define the variable precisely, standardise the measurement method, and train investigators thoroughly before fieldwork begins. Large government surveys in India, such as the Annual Survey of Unincorporated Sector Enterprises, build in frequent data quality reviews during fieldwork precisely so that measurement drift is caught and corrected in near real time rather than discovered after the survey closes.
How people themselves distort measurement
Measurement errors are not always about instruments. They also arise from how respondents report information about themselves. A well-documented example is age heaping: rural respondents, especially those with less formal schooling, tend to round their age to the nearest number ending in zero or five rather than reporting an exact figure. A community survey conducted in villages of the Yavatmal district in Maharashtra found that 42% of respondents reported an age with an incorrect terminal digit, with sharp spikes at ages ending in 0 and 5. Analysis comparing India’s 2001 and 2011 census age data confirms this digit preference is a persistent, if slowly declining, feature of self-reported age across the country.
Income reporting shows a similar, and arguably more consequential, distortion. Households, and particularly wealthier ones, tend to understate their income or consumption, whether out of privacy concerns, tax sensitivity, or simple difficulty in recalling exact figures. This tendency has been one of the biggest obstacles to building a reliable income distribution for India: the government’s new National Household Income Survey was launched specifically because every past attempt to measure income directly collapsed when reported figures came in far below what households were separately reported to be spending and saving. Cross-checking questions, where the same fact is verified through a different angle in the questionnaire, along with intensive investigator training, are the standard remedies for both age and income misreporting.
False data and the problem of bias
A more serious category of non-sampling error is outright fabrication. Field investigators, under time pressure or inadequate supervision, sometimes fill in questionnaires without actually visiting the sampled household or unit, drawing on guesswork or imagination instead of real observation. This is often called cooked-up data, and it is one of the hardest problems to detect after the fact because a fabricated entry can look perfectly plausible on paper. Tight fieldwork supervision, spot checks, and unannounced revisits by senior staff remain the main defence against this problem.
Bias is a subtler distortion. Investigator bias occurs when the person collecting data lets their own assumptions shape what gets recorded. A commonly cited instance is a male head of household answering questions on behalf of the women in the family, filtering their experiences through his own perspective rather than asking them directly. Respondent bias works in the other direction, where the person answering shades their response, consciously or not, based on what they think the interviewer wants to hear, social desirability, or fear of consequences. Both forms of bias are attitudinal rather than technical, which is why training programmes for field staff increasingly focus on sensitising investigators to these blind spots rather than only teaching them how to fill a form correctly.
Recall lapse and problems with the sampling frame
Even an honest, well-trained investigator asking a well-defined question can still get a wrong answer if the respondent simply cannot remember. Recall lapse is the tendency for memory of an event to fade as the reference period lengthens. Ask someone about a health event over the last 15 days and they will likely remember it accurately; ask about the same category of event over the last 365 days and the numbers reported tend to fall, not because fewer events happened but because people forget them. India’s household surveys wrestle with this trade-off directly. For consumption data, the country’s poverty-line methodology has historically used a 30-day recall period for regularly purchased items and a 365-day period only for infrequently bought items, precisely because a shorter window is more reliable for everyday purchases while occasional big-ticket items need a longer window to capture at all. Investigators are also trained to anchor recall to memorable local events, such as a festival or harvest, to help respondents place events in time more accurately.
When the frame itself is faulty
A sampling frame is the complete list of units, households, farms, factories, from which a sample or census draws its units. If this list is outdated, incomplete, or contains duplicate or non-existent entries, every subsequent step of the survey inherits that flaw. A frame built from an old village list will miss new households and may still include ones that have since moved away or merged. There is no statistical fix for a bad frame after data collection; the only genuine solution is careful, up-to-date preparation of the frame before the survey begins.
Non-response and processing errors
Not everyone selected for a survey ends up providing information. Some respondents are simply absent when the investigator calls, others refuse outright, and mail-based enquiries in particular suffer from people not returning forms at all. This is non-response error, and it matters because non-respondents are rarely a random subset of the sample; they often differ systematically from those who do respond. Large-scale surveys such as India’s Household Consumption Expenditure Survey address this by design, with the 2022-23 round visiting each sample household three separate times to complete different parts of the questionnaire, reducing the chance that a single missed visit results in a lost household. Other standard remedies include repeated follow-up requests by mail, personal revisits, drawing an additional sample specifically from the non-responding group, or substituting a non-responding unit with a similar, adjacent unit that was not already part of the sample.
Errors also creep in well after the interview is over. Processing and presentation errors occur during scrutiny of the filled questionnaires, coding of responses into numerical categories, data entry into computer systems, tabulation, and final printing. A single mistyped digit during data entry, or a coding scheme applied inconsistently across batches, can quietly distort results that were collected perfectly well in the field. Quality checks at each processing stage, along with close supervision of coding and entry teams, are the main safeguards here. India’s own experience shows how seriously data quality concerns are taken at this stage: the government’s decision in 2019 to withhold the 2017-18 Consumer Expenditure Survey results over data quality concerns illustrates that even a fully fielded, large-scale national survey can be judged unfit for release if quality checks raise red flags.
Sampling errors: the cost of studying a part, not the whole
All the errors discussed so far, measurement, bias, recall, frame, non-response, and processing, are non-sampling errors: they can occur whether you survey a sample or conduct a full census. Sampling error is different. It exists only because a sample survey studies a subset of the population and then uses that subset to draw conclusions about the entire universe. Since the sample never perfectly mirrors the population in every respect, any estimate drawn from it will differ somewhat from the true value that a complete census would reveal.
The crucial distinction lies in whether this error can be measured. In a non-random sample, where units are chosen through personal judgement or convenience rather than a defined probability method, there is no known mathematical relationship between the sample and the population it is meant to represent. This means the size of the sampling error simply cannot be estimated; you have no way of knowing how far off your sample-based conclusion might be. A random sample, by contrast, gives every unit in the population a known, non-zero chance of selection. This known probability structure makes it possible to estimate the sampling error mathematically and, importantly, to decide in advance how large a sample is needed to keep that error within an acceptable limit. This is precisely why random sampling methods are preferred in rigorous survey design: not because they eliminate error, no sample can do that, but because they let researchers quantify and control it.
Why this distinction matters
Non-sampling errors and sampling errors call for different kinds of vigilance. Non-sampling errors are best tackled through better questionnaire design, thorough investigator training, tight fieldwork supervision, and careful data processing. Sampling error is tackled through the choice of sampling method and sample size. A survey can have a very small sampling error, achieved through a large, well-designed random sample, and still produce misleading results if non-sampling errors such as investigator bias or a faulty frame go unchecked. Conversely, no amount of careful fieldwork can compensate for a poorly designed sample if the goal is to generalise findings to a wider population with a known margin of error. Good data collection practice, whether for a college research project, market research, or a government survey, requires attention to both fronts simultaneously.
What do you think? Between a survey with a small sample but flawless fieldwork, and one with a huge sample but weak investigator training, which would you trust more for a study you cared about? And where in your own experience, filling forms, answering surveys, or even reporting your age, have you noticed people rounding off or adjusting the truth a little?
References
- https://www.mospi.gov.in/uploads/announcements/announcements_1767350115204_48504eeb-1a56-4a89-8077-3affce921293_document_(49).pdf
- https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2963876/
- https://www.sciencedirect.com/science/article/pii/S2213398419303768
- https://www.business-standard.com/economy/news/revamped-survey-looks-to-finally-crack-indian-household-income-puzzle-126051201229_1.html
- https://mospi.gov.in/sites/default/files/reports_and_publication/cso_social_statices_division/Appendix_6_2010.pdf
- https://www.pib.gov.in/PressReleaseIframePage.aspx?PRID=2008737
- https://www.business-standard.com/article/pti-stories/mospi-says-not-releasing-consumer-expenditure-survey-due-to-data-quality-issues-119111501789_1.html
Leave a Reply