Statistics · Psychometrics · Sampling Theory
A Theoretical Justification for the Use of Regression Methods with Psychometric Instruments
The survey case: measurement, sampling, auxiliary variables, and measurement error
An objection occasionally encountered in applied research consists in assuming that regression methods are, for some reason, alien to data obtained through surveys or psychometric instruments. Such a conclusion follows neither from sampling theory nor from statistical theory. A proper answer, however, requires us to distinguish problems that may initially appear identical but in fact belong to different levels of the research process.
The fact that a variable was obtained through a survey does not by itself determine which statistical procedure may be applied to it. What matters is the nature of the variable, the mechanism through which it was measured, the design used to observe it, the properties of the statistical model, and, finally, the inferential problem one is attempting to solve.
The procedure used to collect information and the model used to analyse it are related parts of the scientific process, but they are not the same thing.
1. A survey is not necessarily a psychometric instrument
It is useful to begin with an elementary distinction. A survey is a procedure for obtaining information. It may contain questions about age, income, occupation, number of children, political preferences, consumption, household characteristics, or countless other directly observable or self-reported variables.
A psychometric instrument, by contrast, seeks systematically to measure a characteristic that may not be directly observable: an attitude, ability, trait, perception, or latent construct. It commonly does so through multiple items whose responses are combined in some manner to produce a score.
although a survey may contain one or more psychometric instruments. This distinction matters because problems of sampling, measurement, and statistical modelling are not interchangeable.
2. Regression in Cochran’s theory of surveys
In Sampling Techniques, Cochran explicitly analyses the use of auxiliary information in sample surveys. His starting point is that statistical theory provides very general estimation procedures, while survey practice often benefits especially from methods that can exploit additional information without necessarily requiring a complete specification of the distribution of every attribute under observation.
It is in this context that ratio estimators and regression estimators appear. Suppose that we wish to estimate the population mean of a variable \(Y\) and that an auxiliary variable \(X\), correlated with \(Y\), is available with known population mean \(\bar X\). A classical form of the regression estimator is:
The logic is straightforward. If the auxiliary variable in the sample departs from its known population value, that difference contains information that can be used to adjust the estimate of \(Y\). When the relation between the two variables is sufficiently stable, auxiliary information can increase precision.
Classical sampling theory does not regard regression as incompatible with survey data. On the contrary, it explicitly incorporates estimators based on regression relations in order to exploit auxiliary information and improve estimation of population parameters.
3. Two different uses of regression
It does not follow, however, that every regression model fitted to questionnaire variables is automatically justified by Cochran’s regression estimator. Two different problems must be distinguished:
| Problem | Primary purpose | Example |
|---|---|---|
| Regression estimation in sampling | Use auxiliary information to improve estimation of a population mean, total, or another population parameter. | \(\bar y_{\mathrm{reg}}=\bar y+b(\bar X-\bar x)\) |
| Regression model fitted to survey data | Study associations, make predictions, or represent a conditional relationship between variables. | \(E(Y\mid X)=\beta_0+\beta_1X\) |
Both problems belong to the same statistical tradition and exploit relations among variables, but they do not have exactly the same inferential objective. Cochran’s regression estimator therefore demonstrates clearly that the origin of data in a survey is not an objection to using regression relations; it does not, however, replace the specific justification required for whatever statistical model is subsequently fitted.
4. The specifically psychometric problem
When a variable comes from a psychometric instrument, an additional issue arises: measurement. An observed score should not automatically be identified with the theoretical attribute that the instrument is intended to measure.
In the elementary formulation of classical test theory:
where \(X\) is the observed score, \(T\) the true score in the sense of the model, and \(E\) the measurement-error component.
This has consequences when a psychometric score is used in a regression. Suppose, in the simplest case, that the relationship of interest is:
but that we do not observe \(T\); instead, we observe an imperfect measurement \(X=T+E\). Under the simplest classical model with independent measurement error, using \(X\) as a predictor tends to reduce the estimated slope in absolute value. In this elementary case, attenuation can be expressed approximately as:
The ratio:
corresponds, under this classical formulation, to a notion of reliability. The greater the proportion of variance attributable to measurement error, the lower the reliability and the greater the possible distortion produced by treating the observed score as if it had been measured without error.
With several correlated predictors, the situation may become considerably more complex: measurement error need not result merely in uniform attenuation of all coefficients. For this reason, the psychometric quality of an instrument is not separate from the statistical analysis that follows it.
The fact that a score can mathematically be entered into a regression does not mean that its measurement error has ceased to exist.
5. When regression can be used
There is therefore no general prohibition against applying regression methods to data obtained through psychometric instruments. The correct question is different: which model corresponds to the properties of the variables and to the purpose of the investigation?
If a composite score can reasonably be treated as a quantitative variable and the objective is to represent a conditional mean, linear regression may be appropriate, provided the relevant assumptions are defensible. But the form of the model must change when the structure of the response variable changes.
| Response variable | Possible model |
|---|---|
| Approximately continuous quantitative variable | Linear regression or another appropriate continuous-response model |
| Binary | Logistic or probit regression |
| Ordinal | Ordinal models, such as ordinal logit or probit |
| Count | Poisson, negative binomial, or another count model |
| Latent construct modelled directly | Latent-variable models, SEM, IRT, or other suitable approaches |
It is therefore not the label “psychometric” that decides whether regression can be used. The decision depends on how the variable was constructed, its measurement scale, its measurement error, the form of the response, and the scientific relationship one is attempting to represent.
6. Survey design also matters
One further distinction should not be overlooked. A regression may be correctly specified with respect to its variables and yet produce inadequate inference if the procedure by which observations entered the sample is ignored.
Data from complex sample designs may involve:
- unequal probabilities of selection;
- stratification;
- cluster sampling;
- sampling weights;
- and dependence among observations induced by the design itself.
Under such circumstances, estimation procedures and standard errors may need to account explicitly for the sampling design. The legitimate use of regression on survey data does not imply that the survey design itself can be ignored.
Sampling: how did the units enter the sample?
Measurement: how was the attribute under study transformed into
an observable variable?
Modelling: what probabilistic or functional relationship is
assumed among the variables?
7. A methodological conclusion
A justification for the use of regression methods with survey data does not require claiming that a survey is by nature a psychometric instrument. Nor does it require deducing from regression estimators in sampling theory that every subsequent regression analysis is automatically valid.
The conclusion is more precise. Classical survey theory demonstrates that regression relations can legitimately play a central role even within the sampling-estimation process itself. Psychometrics, in turn, shows that scores obtained through measurement instruments can be analysed statistically, but that reliability and measurement error must form part of the interpretation of the resulting model.
There is therefore no essential incompatibility between surveys, psychometrics, and regression. What exists is a chain of distinct scientific problems, each of which must be solved at its corresponding level:
Each link imposes conditions upon the next. A poorly designed sample is not repaired by using a reliable scale; deficient measurement does not disappear because a regression model has been correctly estimated; and a correctly measured variable does not make correct a model whose mathematical form fails to correspond to the phenomenon.
The origin of a variable in a survey does not by itself determine which statistical method may be applied to it. What matters is the structure of the variable, the process through which it was measured, the sampling design, and the inferential problem one intends to solve.
This is ultimately the general criterion. Statistical methods are not legitimized by the name of the instrument from which the data originate, but by the correspondence between their mathematical properties, the properties of measurement, and the real characteristics of the object under investigation.
References
Cochran, W. G. (1991). Técnicas de muestreo. Mexico: Compañía Editorial Continental.
Fuller, W. A. (1987). Measurement Error Models. New York: John Wiley & Sons.
Lord, F. M., & Novick, M. R. (1968). Statistical Theories of Mental Test Scores. Reading, Massachusetts: Addison-Wesley.
DeVellis, R. F. (2017). Scale Development: Theory and Applications (4th ed.). Los Angeles: Sage.
Lumley, T. (2010). Complex Surveys: A Guide to Analysis Using R. Hoboken, New Jersey: John Wiley & Sons.


Leave a Comment/Deja un Comentario