Topics
A number acquires authority very quickly once it has a label. A device reports a recovery score of 72. A questionnaire produces an anxiety score of 18. An algorithm assigns a risk score of 0.63. A laboratory instrument returns a concentration to three decimal places. Once the value appears on a screen or in a table, it becomes easy to speak as though the number and the quantity named beside it were the same object.
They are not.
The number is an observation produced by a measurement process. The quantity we want to reason about may be only one cause of that observation. Other causes can include calibration, context, behaviour during measurement, properties of the instrument, the wording of questions, environmental conditions, model assumptions, and components that are stable enough to look like signal even when they are irrelevant to the intended interpretation.
The scientific problem is therefore not merely whether a number can be produced consistently. It is whether variation in that number supports the interpretation being assigned to it.
That distinction is old. Cronbach and Meehl's discussion of construct validity in 1955 was built around the problem of interpreting observations as evidence about attributes that are not simply identical to the observations themselves. Modern measurement theory contains several competing accounts of validity, but they share a practical difficulty: the label attached to a variable does not establish what generated its variation.
A simple model makes the problem visible.
An observed score is generated, not discovered
Let (T) denote the quantity we actually want to measure. It may represent a physiological state, a cognitive ability, an environmental property, or another target that is not observed without a measurement procedure.
Suppose the observed score is
[ X = alpha + lambda T + delta B + arepsilon. ]
The constant (alpha) sets the location of the scale. The coefficient (lambda) describes how strongly the target contributes to the score. The variable (B) represents a nuisance component that is not part of the target interpretation but nevertheless affects the measurement. The term (arepsilon) represents occasion-specific noise.
If (delta=0) and (arepsilon) is small, variation in (X) may track (T) closely. If (delta B) is large, the same observed score may primarily reflect something else.
The equation is deliberately generic. In one setting (B) could be a persistent device calibration difference. In another it could represent reading ability contaminating a test intended to measure subject knowledge. In another it could represent a behavioural pattern that changes how a wearable sensor records a physiological signal. The model does not assert that all measurement problems have the same structure. It isolates one fact that is easy to lose once a score has been named: several processes can contribute to the same observable.
This is why measurement is not completed by assigning a unit or a label. The interpretation depends on the mapping between the observable and the target.
A measurement can be almost perfectly repeatable and still measure the wrong thing
Repeatability is valuable. If a supposedly stable quantity produces radically different values under equivalent conditions, something about the measurement process needs explanation.
The converse does not follow. A repeatable measurement need not be a valid measure of the intended target.
Consider two measurements of the same person:
[ X_1 = T + B + arepsilon_1, ]
[ X_2 = T + B + arepsilon_2. ]
Assume that the target (T), the stable nuisance component (B), and the two occasion-specific errors are mutually independent, with
[ operatorname{Var}(T)=1, ]
[ operatorname{Var}(B)=9, ]
and
[ operatorname{Var}(arepsilon_1) = operatorname{Var}(arepsilon_2) = 0.1. ]
Each observed measurement then has variance
[ operatorname{Var}(X_t) = 1+9+0.1 = 10.1. ]
Because (T) and (B) persist across both observations,
[ operatorname{Cov}(X_1,X_2) = operatorname{Var}(T)+operatorname{Var}(B) = 10. ]
The test-retest correlation is therefore
[ operatorname{Corr}(X_1,X_2) = rac{10}{10.1} approx 0.990. ]
By an ordinary repeatability criterion, the measurement looks exceptional.
Now ask a different question. How strongly does the observed score correspond to the target (T)?
Since only (T) is shared between (X) and the target itself,
[ operatorname{Cov}(X,T)=1. ]
The correlation is
[ operatorname{Corr}(X,T) = rac{1}{sqrt{10.1}} approx 0.315. ]
The square of that correlation is approximately (0.099). In this construction, only about ten per cent of the variance in the observed score is linearly associated with variance in the intended target, even though repeated measurements correlate at approximately 0.99.
Nothing is inconsistent about these two results. The stable nuisance component makes the measurement highly repeatable precisely because the nuisance component is stable.
This example should not be interpreted as defining validity by a correlation with a latent variable. Construct validity is a broader question, and different theories of validity formalise it differently. The point is narrower. High reliability cannot by itself establish that a score varies for the reason implied by its interpretation.
The same warning appears in practical measurement guidance. A method can reproduce the same error repeatedly. Precision concerns the dispersion of repeated measurements. Validity concerns whether the interpretation of those measurements is justified.
Random error and systematic contamination behave differently
It is tempting to treat every unwanted contribution to measurement as noise. That language hides an important distinction.
Independent random error can often be reduced by repetition. A stable nuisance component cannot.
Suppose we measure the same target (m) times:
[ X_j = T + B + arepsilon_j, qquad j=1,ldots,m, ]
where the (arepsilon_j) terms are independent and have variance (sigma_arepsilon^2). The average is
[ ar X = T+B+ararepsilon. ]
The variance of the averaged random error is
[ operatorname{Var}(ararepsilon) = rac{sigma_arepsilon^2}{m}. ]
As (m) increases, this term approaches zero. The stable nuisance component (B) does not.
In the limit,
[ ar X longrightarrow T+B, ]
not (T).
This is one reason that collecting more observations can create a false sense of security. Repeated measurement can make an estimate extremely precise around the wrong quantity. The standard error shrinks while the systematic part of the discrepancy remains.
The distinction is not confined to individual instruments. Large observational datasets can estimate biased quantities with very small sampling error when the variables themselves are systematically mismeasured. More data reduce some forms of uncertainty. They do not automatically repair the measurement process that generated the data.
A score can improve while the target remains unchanged
The difference between a score and its target becomes especially important when the score is used to evaluate an intervention.
Suppose a score is defined by
[ X = T + 2C + arepsilon, ]
where (T) is the target construct and (C) is another component that affects the score.
Assume lower values of (X) are interpreted as improvement.
Now introduce an intervention (A). Suppose it has no effect on the target:
[ mathbb E[Tmid A=1] - mathbb E[Tmid A=0] = 0. ]
But suppose it reduces the nuisance component by one unit:
[ mathbb E[Cmid A=1] - mathbb E[Cmid A=0] = -1. ]
If the measurement error has mean zero in both groups, the expected observed change is
[ mathbb E[Xmid A=1] - mathbb E[Xmid A=0] = 0+2(-1) = -2. ]
The score improves by two units while the target does not change at all.
The reverse can also happen. Suppose the intervention genuinely improves the target by two units,
[ Delta T=-2, ]
but increases the nuisance component by one unit,
[ Delta C=1. ]
Then
[ Delta X = -2+2(1) = 0. ]
The target improves while the observed score remains unchanged.
These examples do not imply that composite scores are inherently defective. They show why intervention studies need a theory of measurement as well as a theory of treatment. If an intervention can change components of the measurement process independently of the target, observed score change is not automatically equivalent to target change.
This is particularly important when a score contains behavioural inputs. Once people know how a metric is produced, they may change the inputs that affect the metric without changing the underlying state the metric was intended to summarise. The resulting score is still measured correctly according to its formula. The interpretation is what becomes questionable.
Multiple indicators help only when their errors are informative in different ways
Latent-variable methods often use several indicators rather than a single measurement. This can be a major improvement because different indicators provide partially independent information about a common target.
But the number of indicators is not the decisive property. Their error structure matters.
Consider
[ X_j = lambda_j T + kappa_j C + arepsilon_j, ]
where (T) is the target, (C) is a shared nuisance factor, and (arepsilon_j) is indicator-specific error.
If the indicators have different (lambda_j) values and their nuisance contributions are limited or understood, their joint covariance structure can help identify (T). This is the basic intuition behind many factor models and other latent-variable approaches.
Now consider the simpler case
[ X_j=T+C+arepsilon_j. ]
Averaging many indicators gives
[ ar X = T+C+ararepsilon. ]
Again, the independent errors shrink while the shared nuisance component remains.
A battery of highly correlated measurements can therefore be internally consistent because they all respond to the same unwanted factor. Internal consistency is evidence about a covariance pattern among items. It is not, by itself, evidence that the common source of covariance is the construct named by the scale.
This distinction is one reason modern measurement work treats factor structure, test-retest behaviour, invariance, criterion relations, and substantive theory as separate sources of information rather than expecting one reliability coefficient to settle the interpretation.
Hussey and Hughes illustrated the practical importance of this distinction in a large analysis of commonly used social and personality measures. When scales were assessed mainly through internal consistency, most appeared acceptable. When additional properties such as factor structure, temporal stability, and measurement invariance were considered, the picture changed substantially. Their result concerns a particular collection of psychological measures, not measurement in general, but it demonstrates how conclusions can depend on which properties are examined.
Prediction and measurement are different scientific achievements
A variable can be extremely useful without measuring the construct that people casually say it measures.
Suppose a score (X) predicts an outcome (Y) very well. That predictive relationship may justify using (X) for forecasting under the conditions in which the relationship has been established. It does not automatically justify interpreting (X) as a direct measure of whatever causal quantity is believed to produce (Y).
Prediction asks whether knowing (X) improves our ability to anticipate (Y).
Measurement asks what variation in (X) represents.
Those questions can overlap, but they are not identical.
A proxy can predict because it is downstream of the target, because both share common causes, because it captures a different process that happens to be predictive, or because it encodes features of the data collection system itself. Each possibility can support prediction while implying a different scientific interpretation.
This distinction also explains why replacing an imperfect measurement with a machine learning model does not remove the measurement problem. If the model is trained to reproduce a flawed label, high predictive performance against that label can reproduce the same construct ambiguity at greater scale.
The model may accurately predict the recorded outcome. Whether the recorded outcome represents the scientific target remains a separate question.
The measurement function can change across groups or time
Even a measurement that behaves well in one setting may not support the same interpretation elsewhere.
A simple group-specific measurement model is
[ X_g =
u_g+lambda_gT+arepsilon_g, ]
where (g) indexes a group, context, device version, language, or time period.
If
[
u_1= u_2 ]
and
[ lambda_1=lambda_2, ]
then equal target values generate comparable expected scores under this simplified model.
If either parameter changes, observed score differences become harder to interpret.
Suppose two groups have the same mean value of (T), but the intercept differs by three units:
[
u_2- u_1=3. ]
The second group will have an expected observed score three units higher even though the target distributions are identical.
Or suppose the loading changes. A one-unit difference in (T) may correspond to a one-unit score difference in one setting and a two-unit score difference in another.
These problems motivate measurement invariance analysis. Meredith's formal treatment of factorial invariance made clear that comparisons across populations require assumptions about how observed variables relate to the latent variables being compared.
The same issue arises longitudinally. If people interpret questionnaire items differently after an intervention, if firmware changes the transformation used by a sensor, or if a laboratory assay changes calibration, pre-post score differences may combine target change with measurement-function change.
A numerical difference remains real as a difference in the recorded variable. Its substantive interpretation depends on whether the scale still means the same thing.
More decimal places do not solve a construct problem
A value reported as 72.483 looks more informative than a value reported as 72. That can be true when the extra digits reflect genuine measurement resolution.
But numerical resolution is not the same as epistemic resolution.
If
[ X = T+B+arepsilon, ]
reducing the variance of (arepsilon) improves precision. It does nothing to remove (B).
A device can therefore become increasingly precise while the scientific uncertainty about what the score represents remains largely unchanged. Similarly, a statistical model can produce narrow confidence intervals for an estimand whose connection to the intended construct is poorly defended.
This is not a criticism of precision. Precision is useful because it removes one source of uncertainty and can expose smaller differences. The mistake is to allow precision to substitute for validation.
A narrow interval around the wrong estimand is still an interval around the wrong estimand.
Validation is an empirical programme, not a certificate
The language of a measure being "validated" can suggest that validity is a property awarded once and then carried permanently by an instrument.
That description is too simple for many scientific uses.
Cronbach and Meehl treated construct validation as a process in which an interpretation generates empirical consequences that can be tested. Borsboom, Mellenbergh, and van Heerden later proposed a more explicitly causal account, arguing that a test measures an attribute when the attribute exists and variation in it causally produces variation in the measurement outcome. Other contemporary frameworks differ in important ways, but the common practical lesson is that validity concerns the interpretation and use of measurements, not the visual appearance of the scale or the existence of a reliability coefficient.
Evidence for an interpretation can come from several directions. A measure may behave as expected under controlled interventions, relate to other variables in theoretically constrained ways, separate conditions known to differ in the target, remain invariant across relevant groups, and respond weakly to variables that should be irrelevant.
The pattern matters because different explanations for a score make different predictions.
If a supposed measure of (T) responds strongly when (T) is experimentally manipulated but not when plausible nuisance variables are manipulated, that supports one interpretation. If the score changes more strongly under manipulation of an irrelevant component than under manipulation of the target, that is evidence against the intended interpretation even if the score remains highly reliable.
Good measurement theory therefore creates opportunities for the interpretation to fail.
That is a scientific strength.
Imperfect measurement does not imply arbitrary science
The conclusion should not be that every quantity is hopelessly indirect or that no measurement deserves its name.
Many measurements are extraordinarily successful. Physical instrumentation can be calibrated against standards with well characterised uncertainty. Laboratory assays can be evaluated for specificity, sensitivity, linearity, repeatability, and interference. Psychological and behavioural instruments can be studied through experimental manipulation, latent-variable models, longitudinal data, and cross-population comparisons.
The relevant question is not whether a measurement is perfect.
It is whether the remaining imperfections matter for the inference being made.
A noisy but unbiased measurement may be adequate for estimating a population mean with enough observations. A biased measurement may still rank individuals usefully for a specific prediction problem. A proxy can be entirely appropriate when the scientific claim is explicitly about the proxy. A latent construct can be studied rigorously when the measurement model makes testable commitments about how the construct generates observations.
Problems arise when those distinctions disappear from the language.
A proxy becomes "the thing". A risk score becomes "risk". A questionnaire total becomes "stress". A model output becomes "disease severity". The inferential step from observation to construct is silently removed because the variable name makes the two look identical.
They are not identical.
The quantity of interest belongs in the model
Scientific communication often focuses on uncertainty after the measurement has already been accepted. Confidence intervals, sample sizes, effect estimates, and hypothesis tests receive careful attention.
The measurement process deserves the same scrutiny.
Before asking whether a difference in (X) is statistically significant, we need to know what a difference in (X) means. Before asking whether a model predicts (X) accurately, we need to know why (X) represents the target. Before interpreting change over time, we need evidence that the measurement function has not changed in a way that creates the appearance of target change.
The simple model
[ X=alpha+lambda T+delta B+arepsilon ]
contains the essential warning.
Precision concerns (arepsilon).
Repeatability depends on which components persist.
Validity concerns whether the variation attributed to (T) supports the intended interpretation.
Those properties can move together, but they do not have to.
The most revealing counterexample is the one derived earlier. A score can reproduce itself with a correlation of approximately 0.99 while correlating only approximately 0.31 with the target it is supposed to represent. A stable source of error can look exactly like high-quality signal if repeatability is the only property we inspect.
Measurement becomes scientifically meaningful when the connection between the observation and the target is itself treated as a hypothesis.
The score is evidence.
It is not the thing being measured.
References
Borsboom, D., Mellenbergh, G. J., & van Heerden, J. (2004). The concept of validity. Psychological Review, 111(4), 1061–1071. https://doi.org/10.1037/0033-295X.111.4.1061
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302. https://doi.org/10.1037/h0040957
Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. Advances in Methods and Practices in Psychological Science, 3(4), 456–465. https://doi.org/10.1177/2515245920952393
Fuller, W. A. (1987). Measurement Error Models. Wiley.
Hussey, I., & Hughes, S. (2020). Hidden invalidity among 15 commonly used measures in social and personality psychology. Advances in Methods and Practices in Psychological Science, 3(2), 166–184. https://doi.org/10.1177/2515245919882903
Lord, F. M., & Novick, M. R. (1968). Statistical Theories of Mental Test Scores. Addison-Wesley.
Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance. Psychometrika, 58(4), 525–543. https://doi.org/10.1007/BF02294825
Spearman, C. (1904). The proof and measurement of association between two things. The American Journal of Psychology, 15(1), 72–101. https://doi.org/10.2307/1412159
Embed interactive plots, widgets, and demos using <figure>, <iframe>, or <div class="interactive-embed"> containers. Ensure each embed includes descriptive captions for accessibility.
How to cite
Use the quick export buttons to save citations for reference managers or copy the formatted text directly.
Diogo Ribeiro (2026). Measurement Is Not the Thing Being Measured. Faculty of Media Arts and Design, Technical University of Porto. https://diogoribeiro7.github.io/science-communication/measurement_is_not_the_thing_being_measured/.


