← New search

Other meanings of Survey sampling

STATISTICS

Survey sampling

Survey sampling is the statistical method for selecting samples from populations for surveys. It allows researchers to estimate population characteristics without contacting every member, provided the sample design, selection probabilities, nonresponse, and estimation procedures are appropriate. Probability sampling gives units a known, nonzero chance of selection and supports design-based measures of uncertainty; nonprobability methods can be useful in some settings but require stronger assumptions about representativeness.

Finite population
Target of inference
The complete set of units a survey seeks to describe
Known selection probability
Probability sampling
Basis for design-based estimation
Sampling error
Uncertainty
Variation caused by observing a sample rather than a census
1

Purpose and basic concepts

Survey sampling begins by defining the population, the sampling unit, and the frame from which selections will be made. The target population is the group about which conclusions are intended; the sampling frame is the operational list or other representation used to reach it. A mismatch between the two creates coverage error, such as excluding people without listed telephone numbers or households outside an address frame.1

A census observes every eligible unit, whereas a sample survey observes only a subset. Sampling error is the resulting chance variation, not a mistake in measurement or data processing. Its magnitude depends on sample size, population variability, design, and the estimator. Increasing sample size usually reduces uncertainty, but it does not repair systematic undercoverage or highly selective nonresponse.

Survey estimates commonly describe means, totals, proportions, or relationships. The design must be chosen together with the intended estimates, collection mode, budget, and ethical and privacy requirements.

2

Probability designs

Probability sampling uses a random mechanism so that every eligible unit has a known, nonzero probability of selection. In a simple random sample, units are selected directly from a frame with equal probabilities. Systematic sampling selects units at a fixed interval after a random start, while stratified sampling divides the frame into subgroups and samples within each one.2

Stratification can ensure representation of small or policy-relevant groups and improve precision when strata are internally similar. Cluster sampling selects groups such as schools, villages, or geographic areas rather than individuals; it can reduce field costs, but correlated responses within clusters often increase variance. Multistage designs combine these approaches, for example selecting regions, then households, then people.

Unequal-probability designs are common when some units require oversampling. Their analyses use weights related to the inverse of selection probability, often followed by adjustments for nonresponse and calibration to known population totals.1

3

Estimation and error

Survey estimation must account for the sample design rather than treating observations as an ordinary independent data set. A weighted estimator gives each responding unit an appropriate contribution, while variance estimation uses methods such as Taylor linearization, replication, or specialized design-based software.4

Three broad error sources are distinguished: sampling error, nonsampling error from coverage, nonresponse, measurement, and processing, and model error when estimates rely on assumptions. A reported margin of error generally addresses sampling variability under a specified design and confidence procedure; it does not automatically include all other errors.3

Design effects summarize how a complex design changes variance relative to a simple random sample of the same nominal size. Clustering commonly raises the design effect, whereas effective stratification or carefully chosen unequal probabilities may improve efficiency. Analysts should publish weights, variance methods, response information, and key design decisions so results can be evaluated.

4

Lesser-known aspects

Survey sampling has important edge cases that are easy to miss when only sample size is considered. A large sample from a defective frame may be less informative than a smaller sample with broad coverage. Rare populations may require oversampling, respondent-driven or network designs, or specialized frames, but estimates then depend heavily on weighting and assumptions. Hidden populations, mobile residents, institutionalized people, and people with limited language access can each challenge ordinary household designs.

Nonresponse is not simply a smaller sample: bias occurs when participation is related to the outcome after available adjustments. Follow-up contacts, mixed modes, responsive designs, and calibration can reduce—but cannot guarantee removal of—the problem. Small-area estimation may combine survey data with administrative records or models when direct local samples are too sparse. In online opt-in panels, probability-based recruitment and poststratification can improve inference, but voluntary participation still requires explicit justification rather than an automatic claim of representativeness.5

Glossary

Sampling frame
The list, register, map, or other operational representation from which a sample is selected.
Stratified sampling
A design that divides the frame into strata and selects units separately within them.
Cluster sampling
A design that selects groups of units, such as areas or institutions, before observing elements within them.
Survey weight
A value used to represent the contribution of a sampled or responding unit in estimation.
Design effect
The ratio comparing the variance under a complex design with that of a reference simple random sample.

Margins of error and confidence intervals describe sampling uncertainty under stated assumptions; they are not comprehensive measures of total survey error.