Espartaco

“Is that to say we are against Free Trade? No, we are for Free Trade, because by Free Trade all economical laws, with their most astounding contradictions, will act upon a larger scale, upon the territory of the whole earth; and because from the uniting of all these contradictions in a single group, where they will stand face to face, will result the struggle which will itself eventuate in the emancipation of the proletariat.”

Karl Heinrich Marx · Marx-Engels Collected Works, Vol. VI, p. 290

26,111 views since December 2020

26,111 visitas desde diciembre de 2020

EnglishEspañol

GENERALITIES ON THE STATISTICAL THEORY OF SAMPLING SURVEYS

SURVEY THEORY · SAMPLING · INFERENCE

GENERAL REMARKS ON THE STATISTICAL THEORY OF SAMPLE SURVEYS

Conceptual and technical notes on sampling design, frames, inclusion probabilities, design-based inference, nonresponse, total survey error, and estimation.

The statistical theory of sample surveys is often learned as a collection of procedures: how to select a sample, how to estimate a total or a mean, how to calculate a standard error, and how to determine a sample size. Yet behind every procedure lies a prior conceptual question. What, exactly, is the population to which one intends to infer? What entity is selected at each stage? What makes a selection probabilistic? Where does the randomness that allows us to speak of precision come from? What part of survey error can be quantified by the sampling design, and what part arises from coverage, nonresponse, measurement, or processing?

The notes that follow are organized around problems that appear in the classical sampling literature—in particular Cochran (1977), Kish (1965), and Lohr (2021)—and develop them using tools from contemporary survey theory. The purpose is not to replace the reading of those authors, but to make explicit the conditions under which their statements acquire meaning and the operational consequences that follow from them. When more recent terminology is introduced, the discussion draws especially on Särndal, Swensson, and Wretman (1992), Groves et al. (2009), Biemer and Lyberg (2003), AAPOR’s definitions, and the guidelines of the United Nations, Statistics Canada, and the U.S. Census Bureau.

· · ·

1. In a survey, are costs high and administration more complex?

The answer necessarily depends on the basis of comparison. Relative to a census, a sample survey is usually cheaper and administratively less demanding because it observes only part of the population. Indeed, one of the historical reasons for the development of probability sampling was to obtain sufficiently precise information at a cost far below that of a complete enumeration.

The comparison changes, however, if the reference point is a smaller-scale qualitative procedure, such as a focus group, an in-depth interview, or intensive observation. A national probability survey may require a frame, sample selection, instrument programming, training, supervision, callbacks, weighting, quality control, variance estimation, and documentation. A focus group may be far less expensive, but it does not perform the same inferential function. The choice should not be made by asking which technique is “cheaper” in the abstract, but which technique can answer the research question with the kind of evidence required.

A focus group, for example, can be extraordinarily useful for understanding vocabularies, categories of meaning, mechanisms, and problems in questionnaire comprehension. It is not, however, a direct substitute for a probability survey when the objective is to estimate parameters of a defined population. Cost must therefore be judged in relation to the objective, the required precision, and the type of inference one intends to make.

2. What is a double-barreled question?

It is a question that asks for a single response concerning two or more propositions that could receive different answers. The problem is not that the respondent has “two response options”; almost every closed question does. The problem is that the wording fuses two objects of judgment and forces the respondent to compress them into a single answer.

For example:

Are you satisfied with your salary and with your working conditions?

A person may be satisfied with the salary and dissatisfied with the conditions, or the reverse. If the scale offers only one response, the resulting datum has no univocal interpretation. The solution is to separate the components whenever each constitutes a substantively distinguishable dimension.

This problem points to something more general: a questionnaire is not a neutral container into which already constituted variables are simply “placed.” The concrete wording helps determine the cognitive operation performed by the respondent and, consequently, the datum that is ultimately recorded. Instrument design is therefore part of measurement and not a purely editorial task.

3. When is a questionnaire too long?

There is no universal number of questions or minutes beyond which a questionnaire becomes too long. Burden depends on mode of administration, the cognitive complexity of the questions, the sensitivity of the topics, the population being interviewed, the need to consult records, the number of skip patterns, and the repetition of scales, among other factors.

A face-to-face interview may tolerate a different duration from a telephone interview or a self-administered online questionnaire. Likewise, twenty questions that require recalling amounts, dates, or episodes may impose a greater burden than fifty very brief and homogeneous items.

Pilot testing is important, but it should be distinguished from cognitive pretesting. A pilot reproduces field procedures on a smaller scale and makes it possible to study duration, contact rates, skip patterns, logistics, and implementation problems. A cognitive interview, by contrast, seeks to understand how the respondent interprets the question, retrieves the information, and selects an answer. The U.S. Census Bureau also includes usability testing, behavior coding, debriefing, and split-panel designs among questionnaire evaluation tools.

The methodological rule, therefore, is not “make it as short as possible,” but retain everything needed to measure the defined concepts and remove what adds burden without adding substantive information. Brevity and quality are not synonyms.

4. What does it mean for a frame to contain duplicates?

A sampling frame is the operational device through which sampling units can be identified or reached in order to carry out selection. Sometimes it takes the form of a list; at other times it consists of geographic areas, addresses, administrative registers, potential telephone numbers, or procedures that make it possible to generate selectable units.

Duplication occurs when the same population unit is represented more than once in the selection mechanism. If records are selected with equal probability and one person appears twice, that person may acquire a higher inclusion probability than another who appears only once. But the problem should not be described by saying that any inequality in probabilities makes the sample “nonrandom.” There are deliberately unequal-probability probability designs.

What matters is to distinguish between inequality that is designed and known and multiplicity that is accidental or unknown. A frame may suffer, among other problems, from:

  • undercoverage: units of the target population that cannot be reached through the frame;
  • overcoverage: records that do not belong to the target population;
  • duplicates or multiplicity;
  • outdated records;
  • units whose status or location has changed;
  • many-to-one or one-to-many relationships between records and target units.

The statistical consequence depends on the specific selection mechanism and on whether multiplicity can be measured and corrected in the inclusion probabilities or in weighting.

5. What does data editing mean, and what should be done with nonresponse?

In survey work, “editing” should not be understood as deleting observations that look strange. Editing consists of applying explicit validity, consistency, and plausibility rules to identify records that require review. An improbable answer is not automatically false, and an inconsistent answer should not disappear without a trace.

A reasonable workflow always preserves the raw information and separates the stages:

  1. captured datum;
  2. validation of range, type, and structure;
  3. verification of skips and logical consistency;
  4. possible recontact, when feasible and planned;
  5. documented editing rules;
  6. marking of missing values;
  7. possible imputation;
  8. production of the analytical dataset without destroying the original dataset.

Suppose that a section of the questionnaire applies only to households that reported a certain economic activity. If answers appear in that section for households that, according to the filter question, do not perform that activity, several possibilities exist: the filter was captured incorrectly, the later response was captured incorrectly, the interviewer failed to follow the skip, the definition was misunderstood, or there is a real situation that the design did not anticipate. The datum must be investigated; it is not enough to declare it “obviously wrong.”

It is also necessary to distinguish states that are often collapsed under a single missing-value label: Not applicable, Don’t know, No answer, explicit refusal, failure to contact, structural skip, capture error, and value lost during processing. “Not applicable” is not simply another form of DK/NR: it means that the question does not belong to the logical universe of that unit.

Unit nonresponse and item nonresponse

Unit nonresponse occurs when no substantive information is obtained from a selected unit; item nonresponse occurs when an interview exists but particular answers are missing. The former is often addressed through weighting adjustments and the latter may require imputation or other methods, although the specific techniques depend on the missingness mechanism and the analytical objective (Kalton and Kasprzyk, 1986).

A high nonresponse rate does not mechanically imply that a variable is unusable, just as a high response rate does not demonstrate the absence of bias. Nonresponse bias depends on the relationship between response propensity and the variables to be estimated.

AAPOR distinguishes, among other measures, response rate, cooperation rate, contact rate, and refusal rate. They are not synonyms. The response rate relates completed interviews to the set of eligible cases—with variants concerning the treatment of unknown eligibility and partial interviews; cooperation conditions on contacted units; the contact rate studies the success of contact; and the refusal rate quantifies a specific form of nonresponse.

Finally, interviewer probing must be standardized. The objective is not to press until “no survey is ever left incomplete,” but to facilitate a valid answer without inducing it and while respecting the participant’s consent.

6. Efficiency, repeated sampling, and confidence intervals

Cochran frames sampling theory as a joint problem of precision and cost: to design selection and estimation procedures that produce the required accuracy at the lowest possible cost. The underlying idea remains fundamental. There is no “best” design outside an objective; there is a design that is more or less efficient for particular estimands, constraints, and costs.

The precision of a procedure is studied through its distribution under hypothetical repetitions of the sampling mechanism. From a design-based perspective, the values of the finite population may be regarded as fixed, while randomness arises from the set of samples that the design could have selected.

Here it is important to state the interpretation of a confidence interval precisely. If a procedure constructs intervals with a 95% confidence level, that does not mean that 95 out of 100 samples will produce “values similar” to the value observed in our sample. It means that, under the assumptions of the procedure and under hypothetical repetitions of the design, approximately 95% of the intervals constructed in that way would contain the fixed population parameter.

Once a particular frequentist interval has been observed, the parameter does not thereby become random. The randomness belongs to the procedure that would have produced different intervals under other possible samples.

Sample size for a proportion

In an elementary formulation under simple random sampling, if an approximate margin of error \(e\) is desired for a proportion at a level associated with \(z_{1-\alpha/2}\), one may begin with

\[ n_0=\frac{z_{1-\alpha/2}^2\,p(1-p)}{e^2}. \]

When there is no useful prior information about \(p\) and a conservative case is sought, \(p=0.5\) is used because the function

\[ p(1-p) \]

attains its maximum at \(p=0.5\). That is the mathematical reason. It is not a consequence of the central limit theorem.

If the population is finite and the sampling fraction is no longer negligible, sample size can be adjusted, within the same simple framework, by

\[ n=\frac{Nn_0}{N+n_0-1}. \]

In complex designs, the calculation must incorporate stratification, clustering, unequal probabilities, domain objectives, anticipated nonresponse, and, where appropriate, an expected design effect. Sample size should not be calculated in isolation from the design.

7. Are there “common sample sizes”?

Recurring sizes—400, 800, 1,000 interviews, for example—do appear in professional practice because certain combinations of precision, budget, and mode recur. But those numbers are not statistical constants.

For a proportion near 0.5, a simple random sample of about 1,067 units yields, under a normal approximation and without finite population correction, a sampling margin close to three percentage points at 95%. This accounts for part of the recurrence of sample sizes near one thousand in polls. The calculation, however, does not by itself demonstrate that a survey is nationally representative or that its total error is ±3 percentage points.

If there is clustering, highly variable weights, small domains, unequal response rates, or incomplete coverage, the same \(n\) may yield very different precision and quality. The proper question is not “what is the traditional size?” but “what precision do I need for which estimand, under what design, and at what cost?”

8. Two differences between classical theory and survey theory

8.1. Distributions, models, and design

A classical observation by Cochran contrasts part of model-based statistical inference with survey theory for finite populations. In many classical problems, observations are assumed to arise from a distribution with a specified mathematical form and its parameters are estimated. In survey sampling, by contrast, a central tradition develops inference from the selection mechanism without requiring a parametric distribution for the population values.

The expression “few measurements on each unit” does not refer to a sample with fewer than 30 or 50 cases. It refers to the number and type of variables measured on each unit. A large survey may collect many variables with very different distributions; there is no reason to impose the same parametric family on all of them merely because the dataset is large.

Computing power does not eliminate this conceptual difference. Today at least three perspectives coexist:

  • design-based inference: randomness is attributed to the selection mechanism;
  • model-assisted inference: models are used to improve estimators, calibrate, or exploit auxiliary information while retaining properties justified by the design;
  • model-based inference: inference rests more directly on a probability model for the observed values or mechanisms.

Särndal, Swensson, and Wretman (1992) systematically develop the model-assisted perspective. The distinction is not a relic from before modern computers; it answers different questions about the source of randomness and the assumptions that support inference.

8.2. Finite populations and the finite population correction

A population is finite if it contains \(N<\infty\) units. This does not depend on whether the researcher knows \(N\) exactly. A finite population can exist even when its size is unknown or uncertain.

Under simple random sampling without replacement, for the population mean \(\bar Y\), if

\[ S^2=\frac{1}{N-1}\sum_{i=1}^{N}(Y_i-\bar Y)^2, \]

then

\[ \operatorname{Var}(\bar y) =\left(1-\frac{n}{N}\right)\frac{S^2}{n}. \]

The factor

\[ 1-f,\qquad f=\frac{n}{N}, \]

is the finite population correction in the variance. When \(f\) is very small, \(1-f\approx1\) and its effect may be negligible. When the sample covers a large proportion of the population, ignoring it can substantially overestimate sampling variance.

The reason is materially simple: if an increasing proportion of a finite population is observed without replacement, progressively less uncertainty remains about the part that was not observed. When \(n=N\), a census of that population has been conducted and the variance due to the sampling mechanism is zero, although measurement, coverage, processing, and other errors may still exist.

9. What does it mean to be sensitive to respondents’ expectations?

The interviewer must maintain a combination that is not always easy: rapport without leading. The interviewer should establish a relationship that facilitates communication, explain the purpose of the interview, listen, and use permitted probes; but should not guide the respondent toward an answer that appears socially desirable or compatible with the researcher’s expectations.

Empathy does not mean abandoning standardization. An interview can be humane, respectful, and sensitive to context without arbitrarily altering the meaning of the questions. In standardized instruments, an important part of training consists precisely in learning which clarifications are permitted and how to use neutral probes.

Sensitivity also has an ethical dimension. The interviewer must recognize discomfort, fatigue, privacy, risks, and the right not to answer. Data quality cannot be pursued through coercion.

The interaction between interviewer and respondent can affect measurement. Tone of voice, the way response options are presented, verbal or gestural reactions, the respondent’s perception of the interviewer’s status, and topic sensitivity can alter the observed response.

Several mechanisms should be distinguished:

  • social desirability: a tendency to provide an answer perceived as socially acceptable;
  • acquiescence: a tendency to agree with formulations regardless of content;
  • interviewer effect: systematic variation associated with interviewer characteristics or behavior;
  • mode effect: differences associated with face-to-face, telephone, web, paper, or other modes;
  • context and order effect: responses conditioned by preceding questions or by the structure of the instrument.

Informed consent, however, is not a bias. It is an ethical condition of research involving human participants. A different question is whether the decision to participate may generate self-selection or differential nonresponse. That possible source of bias should be studied as part of the participation process, not presented as a reason to weaken consent.

11. Survey design and absolute sample size

Lohr insists on an idea that should be formulated carefully: an enormous sample does not rescue a defective design. The relevant quantity in her discussion is the absolute size of the sample, not the absolute size of the population.

For large populations, the precision of an estimator under simple random sampling depends primarily on \(n\); \(N\) appears in the finite population correction and has little effect when \(n/N\) is small. Thus, a sample of one thousand people may have a similar sampling error whether the population contains one million or one hundred million people, provided that the selection mechanism and other assumptions are comparable.

But this statement does not mean that any sample of one thousand units is adequate. A convenience sample of one hundred thousand people may have more selection bias than a probability sample of one thousand. Increasing \(n\) reduces random variability under the assumed mechanism; it does not automatically erase systematic errors of coverage, selection, measurement, or nonresponse.

Hence a rule that should remain present whenever polls are read:

\[ \text{large sample}\;\nRightarrow\;\text{representativeness}. \]

And likewise:

\[ \text{small sampling margin}\;\nRightarrow\;\text{small total error}. \]

12. Content, units, extent, and time when defining a population

A study population must be defined precisely enough to decide who or what belongs to it. It is useful to specify its substantive content, the units that compose it, its spatial extent, and the time or reference period.

Content is not simply spatiotemporal location; that belongs mainly to extent and time. Content determines what kind of entity and what substantive condition make a unit part of the population. For example: “persons aged 18 or older who are usual residents of private households in Costa Rica during the reference period.” There we find an entity, an age condition, a residence condition, a household delimitation, and a territorial-temporal extent.

The different units

The word “unit” has a general meaning—something that may be considered an entity for a given analysis—but in surveys it acquires different technical functions. They should not be confused.

Element or target unit. The population entity about which information is to be produced: a person, a household, a firm, a plot of land, a school.

Unit of analysis or study unit. The entity to which the analyzed variables are attributed. It may coincide with the target element, though not always.

Observation unit. The entity on which observation or measurement is directly performed.

Reporting unit. The person or entity that provides the information. In a household survey, one adult may report on other members; the reporting unit and the analytical unit are then not identical.

Sampling unit. A selectable unit at a particular stage of the design. At a first stage it may be a geographic segment; at a second, a dwelling; at a third, a person.

Primary sampling unit (PSU). The unit selected at the first stage of a multistage design. It is not defined by being “a place” or by necessarily containing only one final unit.

Listing unit. A unit identified during a listing operation in order to construct or update the frame for a later stage. Its function is operational.

The same entity may perform several functions. If a complete list of students is available and students are directly selected and asked about themselves, a student may simultaneously be the element, unit of analysis, observation unit, reporting unit, and sampling unit. But this empirical coincidence does not make the concepts synonymous.

13. Domain, stratum, class, and subclass

An estimation domain is a subpopulation for which separate estimates are desired. It may be defined by region, sex, age, economic sector, or another relevant condition.

A stratum, by contrast, is a partition used by the sampling design before the sample is selected. Strata are generally intended to control composition, improve efficiency, or ensure sufficient representation of relevant groups.

A domain may coincide with a stratum, but it need not. It may cut across several strata or be defined using information obtained after selection. For example, the population may be stratified by province and, later, a domain may be estimated consisting of persons employed in a particular economic industry across all provinces.

The notion of a class as an interval in a frequency distribution belongs to a different problem. It is not necessary for defining an estimation domain. Likewise, “subclass” should not be used as a general synonym for domain unless a source explicitly defines that terminology for a particular scheme.

14. What does \(Y_i\) represent?

If the finite population is

\[ U=\{1,2,\ldots,N\}, \]

and \(Y\) is a study variable, then \(Y_i\) represents the value taken by that variable for unit \(i\). It does not represent “the characteristics” of the unit in the plural, but one particular realization of a defined variable.

The population total is

\[ T_Y=\sum_{i=1}^{N}Y_i, \]

and the population mean is

\[ \bar Y=\frac{1}{N}\sum_{i=1}^{N}Y_i. \]

In design-based inference, a particularly important idea is that the vector

\[ (Y_1,\ldots,Y_N) \]

may be treated as fixed. The distribution used to study an estimator arises from the sampling design: which samples \(s\) could have been selected and with what probabilities.

This prevents a common confusion. The use of probability does not require us to think that every material characteristic of the population was “generated” by a known parametric distribution. Probability may reside in the act of selection.

15. What does MESIP selection mean?

In terminology used in various Latin American texts, MESIP commonly denotes selection with equal probability of the elements. It should be understood as a particular class of probability design, not as the definition of all probability sampling.

If \(\pi_i\) is the inclusion probability of unit \(i\),

\[ \pi_i=P(i\in s), \]

an equal-probability design satisfies, for comparable units within the design,

\[ \pi_i=\pi_j. \]

But a general probability design may have

\[ \pi_i\neq\pi_j. \]

The inequality may be deliberate, for example when selection is made with probability proportional to a measure of size.

16. What does “unrestricted random sampling” mean?

The expression is used in part of the Spanish-language literature to refer to simple random sampling without replacement: from a population of size \(N\), a fixed-size sample \(n\) is selected so that every possible subset of \(n\) units has the same probability of selection.

If \(S\) is the set of all possible samples of size \(n\), then for every \(s\in S\),

\[ P(S=s)=\binom{N}{n}^{-1}. \]

It follows that every unit has inclusion probability

\[ \pi_i=\frac{n}{N}. \]

It is a fundamental design precisely because it permits transparent derivations and serves as a benchmark. But it is not “random sampling as such” in an exhaustive sense: probability sampling includes stratified, systematic, cluster, multistage, and unequal-probability designs.

17. Mean squared error, variance, and bias

Let \(\hat\theta\) be an estimator of the parameter \(\theta\). Its mean squared error is

\[ \operatorname{MSE}(\hat\theta) =E\left[(\hat\theta-\theta)^2\right]. \]

It can be decomposed as

\[ \operatorname{MSE}(\hat\theta) =\operatorname{Var}(\hat\theta) +\operatorname{Bias}(\hat\theta)^2, \]

where

\[ \operatorname{Bias}(\hat\theta) =E(\hat\theta)-\theta. \]

The identity can be seen by writing

\[ \hat\theta-\theta =\bigl(\hat\theta-E(\hat\theta)\bigr) +\bigl(E(\hat\theta)-\theta\bigr), \]

squaring, and taking expectations. The cross term vanishes because

\[ E\left[\hat\theta-E(\hat\theta)\right]=0. \]

Bias is not the difference between one observed realization of \(\hat\theta\) and its expectation. That difference is a random deviation of the realization from the center of its distribution. Bias is a property of the estimation procedure: the difference between the estimator’s expectation and the parameter it is intended to estimate.

The decomposition is conceptually important because it displays a possible trade-off between variability and bias. An estimator with slightly larger variance may have smaller MSE if it reduces bias enough, and vice versa. “Precision” and “accuracy” should not be collapsed into a single colloquial intuition when the analysis requires their components to be distinguished.

A cluster in sampling is a grouping of elementary units that may be used as a selection unit. It should not be confused with a cluster produced by a machine-learning algorithm. In clustering, the grouping is usually the result of an algorithmic criterion of similarity or distance; in sampling, the cluster belongs to the design and may be defined by geography, administrative organization, dwellings, schools, establishments, or other structures.

A PSU is simply the unit selected at the first stage. In many household designs the PSU is a geographic cluster—for example, a census segment—so the two notions coincide empirically. They are not, however, synonymous by definition.

Design First unit selected Stages Are equal probabilities mandatory?
Simple random Element 1 Yes, in the classical simple design
Stratified Element within stratum 1 Not necessarily across strata
One-stage cluster Cluster 1 No
Multistage PSU; then secondary units, etc. 2 or more No
PPS Unit with probability proportional to size 1 or more No

Clustering matters because units within the same cluster often resemble one another. If there is positive intraclass correlation, interviewing twenty people within the same cluster may provide less information than interviewing twenty people spread across many clusters.

This relative loss or gain in precision is often summarized through a design effect,

\[ \operatorname{deff} =\frac{\operatorname{Var}_{\text{design}}(\hat\theta)}{\operatorname{Var}_{\text{SRS}}(\hat\theta)}, \]

which compares the variance under the actual design with a simple-random-sampling reference of comparable size. A \(\operatorname{deff}>1\) indicates greater variance than the reference; it is not an “error” of an algorithm, but a property that may arise from the design, intraclass correlation, and weight variability.

19. What does “margin of error” mean?

In the language of polls, the margin of error is usually the half-width of a confidence interval associated with sampling uncertainty under a particular design and estimation method. It is not “the total amount of error in the survey.”

For an estimate \(\hat\theta\), a generic form is

\[ \hat\theta\pm z_{1-\alpha/2}\,\widehat{SE}(\hat\theta), \]

so that the margin is

\[ MOE=z_{1-\alpha/2}\,\widehat{SE}(\hat\theta). \]

In complex designs, \(\widehat{SE}\) must correspond to the design: strata, clusters, unequal probabilities, calibration, or replicate weights, as applicable. Calculating a margin as though the sample had been simple random when it was not can understate or overstate uncertainty.

Sampling margin and total survey error

The Total Survey Error paradigm requires separating sampling variability from other sources of error. These include:

  • coverage error;
  • nonresponse error;
  • measurement error;
  • construct or specification error;
  • processing and coding error;
  • editing and imputation error;
  • frame and linkage error;
  • errors arising from models or adjustments.

Thus a phrase such as “the survey has a margin of error of ±3%” should be understood, unless otherwise specified, as a statement about one component of uncertainty. It does not demonstrate that the true value lies within three points or that every other source of error is small.

20. Why is probability sampling preferred? Can only probability samples be useful?

A probability design provides a known selection distribution and therefore permits design-based inferences and estimation of sampling variability under explicit conditions. This is an enormous advantage when the objective is to infer to a defined population.

But “nonprobability” does not automatically mean “useless.” Convenience, judgment, quota, network, or self-selected samples may be appropriate for exploration, qualitative research, instrument development, mechanism studies, hard-to-reach populations, or questions for which design-based population inference is not the objective.

The problem arises when such samples are attributed properties that their selection mechanism does not justify. If a researcher studies a localized group because it possesses a rare characteristic, valid knowledge may be produced about that group or about particular mechanisms. One cannot infer without additional assumptions that the prevalence of that characteristic in the country is equal to the prevalence observed in the group.

The word representativeness should be used cautiously. A probability sample may suffer from undercoverage, differential nonresponse, or measurement error. It is therefore preferable to state exactly what property has been secured: a defined coverage frame, known probabilities, calibration to particular totals, balance on auxiliary variables, or inferential capacity under a specified design.

Methodological ethics require documenting the selection method and the limits of inference. A nonprobability sample should not be presented as probabilistic; a probability sample should likewise not be presented as free from the remaining sources of survey error.

21. “Randomization is a physical operation”: what does that mean?

Kish insists that a probability sample requires more than writing a distribution on paper. There must be an effective selection procedure whose operation is congruent with the specified probability model.

The word “physical” should not be understood as opposed to computation. An algorithm implemented on a computer can perfectly well form part of the material operation of selection. What Kish contrasts is an actually executed chance mechanism with a mere belief, assumption, or declaration that the units “are as if they had been random.”

We can separate two levels:

\[ \text{probability model}\neq\text{randomization procedure}, \]

but a probability design requires correspondence between them. If the protocol states that each of \(N\) units has equal probability of selection, the actual procedure—a random-number table, an audited pseudorandom algorithm, a draw, or another mechanism—must implement that property.

This distinction is philosophically interesting. Probability in a probability survey is not merely a formal property of the language used to describe selection; it is linked to an objective, reproducible practice. The theory acquires inferential capacity because a controlled relation exists between the material mechanism and the mathematically defined sample space.

22. What does it mean that the frame includes “physical lists” and procedures that do not literally list all units?

Kish uses frame in a broader sense than “a file containing every name.” A physical list may be a register of persons, establishments, addresses, plots, or administrative units. But there may also be procedures that allow units to be selected without constructing an exhaustive list of all final elements.

An area frame, for example, may divide a territory into selectable segments; dwellings are then listed within the selected segments. In random-digit dialing, the frame may be conceptual or algorithmic: the space of eligible numbers is defined and selections are generated without previously possessing a file of all subscribers.

“Actually listing” means explicitly enumerating all relevant units in an operational list. Avoiding that effort does not mean replacing the physical by “the computational” in the abstract; it means constructing a coverage and selection mechanism that does not require an exhaustive prior enumeration of all final units.

23. Why is it so difficult to construct a frame for human populations?

Because target population, frame, and field location are not the same thing. An electoral register, for example, may be excellent for a population of registered voters and at the same time inadequate for a population defined as “all resident persons” if it excludes minors, unregistered persons, or other groups.

Human populations also change. People are born, die, migrate, move homes, form or dissolve households, use multiple telephones, lose numbers, share devices, and may reside temporarily in different places. A perfect frame would be costly not because we need to know where every person is at every hour, but because it must maintain a sufficiently good correspondence between target units and selectable units.

Large statistical systems address this problem through different strategies: address frames, integrated administrative registers, area frames, dwelling listings within PSUs, dual or multiple frames, and periodic updates. Each solution distributes costs and risks of undercoverage, duplication, and obsolescence differently.

The correct methodological question is not “does a complete list of every inhabitant exist?” but “what frame or combination of frames provides a documented path by which units of the target population have inclusion probabilities that are known or can be modeled in accordance with the design?”

24. What is multistage sampling?

It is a design in which selection occurs in two or more stages. In a household survey, for example, one might select:

  1. PSUs or geographic segments;
  2. dwellings within the PSUs;
  3. one person within each dwelling.

If the selections are conditional, the overall inclusion probability of person \(i\) may be written schematically as the product of the probabilities at each stage:

\[ \pi_i =\pi_{h}\,\pi_{j\mid h}\,\pi_{i\mid j,h}, \]

depending on the notation adopted for PSU, dwelling, and person.

The basic design weight is then

\[ w_i=\frac{1}{\pi_i}. \]

In practice, that weight may later be modified for nonresponse, calibration, poststratification, weight trimming, or other procedures. The basic design weight must always be distinguished from final analytical weights.

Multistage sampling makes it possible to work with enormous populations without first constructing a national list of every final element. Its cost is usually much lower, although geographic concentration can increase variance relative to a simple random sample of the same size.

25. Multiplicity, unequal probabilities, and Horvitz-Thompson estimation

Suppose a telephone frame allows a household with two lines to appear twice while another household with one line appears once. If lines are selected with equal probability and households are the final target, household inclusion probabilities may be unequal. If that multiplicity is unknown and ignored, it may produce bias.

But it does not follow that all random sampling requires equal probability. Survey theory explicitly includes unequal-probability designs. What is fundamental is that the probabilistic mechanism be specified and that the probabilities required for estimation be known or calculable.

For a population \(U\) and a total

\[ T_Y=\sum_{i\in U}Y_i, \]

if each unit has positive inclusion probability \(\pi_i>0\), the Horvitz-Thompson estimator is

\[ \widehat T_{HT} =\sum_{i\in s}\frac{Y_i}{\pi_i}. \]

Its logic is immediate: a unit that is difficult to select represents more units of the population than a unit with a higher inclusion probability. Hence the inverse of \(\pi_i\) appears.

Under the appropriate design,

\[ E_p\left(\widehat T_{HT}\right)=T_Y, \]

where \(E_p\) denotes expectation with respect to the sampling design. The historical importance of the Horvitz and Thompson (1952) result lies precisely in showing how samples without replacement and with unequal probabilities can be handled without thereby turning them into “nonrandom” samples.

This allows us to return to the problem of duplicates with greater precision. If multiplicity changes \(\pi_i\) in a known way and the estimator incorporates it correctly, the inequality can be handled. If multiplicity is unknown, it becomes a frame problem that prevents the probabilities from being properly known and can distort inference.

· · ·

Some General Consequences

Four principles that cut across survey theory can be drawn from the preceding problems.

First, design and estimation form a unity. There is no estimator whose inferential meaning can be completely separated from the way the sample was produced. Weighting, variance estimation, and the interpretation of intervals depend on the design.

Second, randomness does not necessarily mean equality. A design may be probabilistic and assign unequal probabilities; what matters is that the mechanism be probabilistically defined and that its relevant probabilities can be incorporated into inference.

Third, sampling error is not total survey error. A narrow interval may coexist with undercoverage, differential nonresponse, poor measurement, or processing errors. The mathematical precision of one component should not be confused with overall validity.

Fourth, the statistical operation is also a material operation. Frames, listings, interviews, selection algorithms, questionnaire skips, and editing rules are concrete practices. Theory does not float above them: its assumptions are realized—or violated—in the actual process by which data are produced.

References

AAPOR. (2023). Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys (10th ed.). American Association for Public Opinion Research. https://aapor.org/standards-and-ethics/standard-definitions/

Biemer, P. P., & Lyberg, L. E. (2003). Introduction to Survey Quality. Wiley. https://doi.org/10.1002/0471458740

Cochran, W. G. (1977). Sampling Techniques (3rd ed.). Wiley.

Groves, R. M. (1989). Survey Errors and Survey Costs. Wiley. https://doi.org/10.1002/0471725277

Groves, R. M., Fowler, F. J. Jr., Couper, M. P., Lepkowski, J. M., Singer, E., & Tourangeau, R. (2009). Survey Methodology (2nd ed.). Wiley.

Horvitz, D. G., & Thompson, D. J. (1952). A generalization of sampling without replacement from a finite universe. Journal of the American Statistical Association, 47(260), 663-685. https://doi.org/10.1080/01621459.1952.10483446

Kalton, G., & Kasprzyk, D. (1986). The treatment of missing survey data. Survey Methodology, 12(1), 1-16.

Kish, L. (1965). Survey Sampling. Wiley.

Lohr, S. L. (2021). Sampling: Design and Analysis (3rd ed.). CRC Press.

Särndal, C.-E., Swensson, B., & Wretman, J. (1992). Model Assisted Survey Sampling. Springer.

Statistics Canada. (2003). Survey Methods and Practices. Catalogue no. 12-587-X. https://www150.statcan.gc.ca/n1/en/pub/12-587-x/12-587-x2003001-eng.pdf

United Nations Statistics Division. (2008). Designing Household Survey Samples: Practical Guidelines. Studies in Methods, Series F, No. 98. https://unstats.un.org/unsd/publication/seriesf/seriesf_98e.pdf

U.S. Census Bureau. (2021). Questionnaire Testing and Evaluation Methods for Censuses and Surveys. https://www.census.gov/about/policies/quality/standards/appendixa2.html

Additional Resources of Interest

The following resources may be useful for quick orientation, discussion, examples, or bibliographic navigation. They do not replace the methodological sources cited above.


Discover more from Marxist Philosophy of Science

Subscribe to get the latest posts sent to your email.

Follow the blogSeguí al blog

Comments

Leave a Comment/Deja un Comentario

Discover more from Marxist Philosophy of Science

Subscribe now to keep reading and get access to the full archive.

Continue reading