Causal inference
Measurement Error, Proxies, and Misclassification
Analyze treatment, outcome, and confounder measurement error, differential misclassification, proxy variables, and validation designs.
By the end you can
- Distinguish construct, operational definition, and recorded measurement
- Explain differential and nondifferential measurement error
- Recognize proxy confounding and treatment misclassification
- Design validation, repeated-measure, or sensitivity analyses
Example
A prescription record mistaken for medication use
A study needed to know who had been treated, and what it had in the records was prescription orders. So it used them. Order present, patient counted as treated; order absent, patient counted as untreated.
The rule is easy to apply and it fails in three ordinary ways at once. Some patients never filled the prescription and were counted as treated anyway. Others obtained the medication elsewhere and sat quietly in the untreated group. Among those who did fill it, adherence varied enough that one label covered a full course and a box opened once.
So the column recorded a clinician's intention. The write-up discussed a drug acting in a body.
Nobody decided to change the question the study was answering. The recording rule changed it for them, and the quantity finally estimated was not the quantity the discussion section argued about.
Every other link in the chain has its own version of this. A billing code tells you something was reimbursed, which is not the same as delivered. An outcome only becomes an event when somebody looks for it, so if one arm is followed more closely than the other, the counts diverge before the disease does. A severity proxy captures what was written down and misses the clinician's unwritten worry that drove the decision to treat. And consent, which decides who is in the study at all, is often collected during a visit that the treatment itself caused.
None of those links is a thought experiment. Each has been measured, in named studies, on hundreds of thousands of patients, and the rest of this lesson works through the measurements rather than the worry.
- The construct is the thing the study is really about — the medication acting in the body, not any record that it was ordered.
- The operational definition is the rule that turns that construct into a value in a column, and it is written by people who have to work with the data that happens to exist.
- Misclassification is what that rule produces when it puts someone in the wrong category: treated when they never took anything, untreated when they took it from somewhere else.
- A validation sample is a subset measured a better way, and it is the only thing that turns a suspicion about the gap between rule and construct into a number.
Causal variables are measurement systems, not column names
The prescription column is not an unlucky special case. Treatments, outcomes and confounders all arrive as something else — a code, a sensor reading, a survey answer, a clinician's judgment, or a frank stand-in for a quantity nobody can observe at all. Each of those is a small measurement system with its own way of being wrong.
Being wrong does not have one direction. Depending on how the error is structured it can pull an effect toward nothing, push it further from nothing, reverse its sign, or produce an association where the truth is a flat line. Two studies can misclassify the same share of people and end up with opposite problems. What matters is not how often the instrument errs but what makes it err.
The tolerable kind is error with no connection to anything under study. Noise that would have fallen the same way whichever arm the patient was in, at least under conditions you can state out loud. It usually blurs an effect rather than inventing one. The blur has been sized.
Nine prospective observational studies all recorded each person's blood pressure the same cheap way: a single baseline diastolic reading. Between them they followed 420,000 individuals for 6-25 years, a mean of 10, and counted 843 strokes and 4,856 coronary events. One morning's reading is not a person's usual pressure. The gap between them is noise of exactly the harmless-looking sort. Corrected for that regression dilution, prolonged differences in usual DBP of 5, 7.5 and 10 mm Hg were associated with at least 34%, 46% and 56% less stroke, and 21%, 29% and 37% less coronary heart disease. MacMahon and colleagues pooled the studies for The Lancet in 1990, and the abstract states the size of what the recording rule had hidden: “These associations are about 60% greater than in previous uncorrected analyses.” The effect had always been that big. The column had been reporting a shrunken copy of it.
The dangerous kind depends on the treatment, on the outcome, or on how people entered the sample at all. That error is already correlated with something under study before any model is fitted.
A pulse oximeter is the cleanest demonstration that precision and dependence are different properties. Oximeter readings were paired with arterial blood gas co-oximetry taken within 10 minutes: 10,789 pairs from 1,333 White and 276 Black patients at the University of Michigan, and 37,308 pairs from 7,342 White and 1,050 Black patients in intensive care units at 178 hospitals. Among readings the device placed at 92-96%, arterial saturation was actually below 88% in 11.7% of measurements in Black patients against 3.6% in White patients at Michigan, and 17.0% against 6.2% in the multicentre cohort. Sjoding and colleagues published the comparison in the New England Journal of Medicine in 2020, and stated the result plainly: “Thus, in two large cohorts, Black patients had nearly three times the frequency of occult hypoxemia that was not detected by pulse oximetry as White patients.”
A separate Johns Hopkins study looked at 7,126 COVID-19 patients. Among the 1,216 who had concurrent arterial and oximeter measurements, occult hypoxemia turned up in 28.5% of Black against 17.2% of White patients. Among a further 1,903 patients, Black patients had a 29% lower hazard of having treatment eligibility recognised (HR 0.71, 95% CI 0.63-0.80). That is the point at which a measurement error stops being a measurement error. It has become a treatment assignment.
Confounders are the quietest of the three failures. A proxy stands in for something latent — concern, severity, frailty — and the analyst adjusts for the proxy because the proxy is what exists. That adjustment removes the part of the confounding the proxy managed to capture. The rest stays inside the estimate, where it is indistinguishable from effect.
An effect can be estimated to the third decimal and still belong to a variable nobody meant to study.
Analogy
One thermostat, and every room in the building called comfortable
One thermostat on a corridor wall reports the temperature of a whole building. It reports its own square of wall very well — to a fraction of a degree, repeatably, all day. What it cannot report is the room on the cold side, the draft under the loading door, or the fact that it was mounted beside a radiator by whoever was holding the drill.
The reading is precise. How warm it is in here, for the people in here, is only loosely related to it, and adding decimal places to the sensor does nothing to close that distance.
A wall sensor does have one mercy: it does not care what anyone decides. It will not drift toward the boiler because the heating was switched on. Causal measurements do precisely that. Billing codes are generated by the act of treating, follow-up visits are scheduled because of it, consent is asked for during it. The instrument moves with the treatment, and that is the whole difference between a reading that is merely coarse and one that is bent toward an answer.
How far it bends has been measured on close to a million patients. Hospitals differ enormously in how hard they look for venous thromboembolism: 32 diagnostic imaging studies per 1,000 in the lowest-use quartile, 167 per 1,000 in the highest. Risk-adjusted VTE event rates climbed stepwise across those same quartiles, from 5.0 per 1,000 to 13.5 per 1,000. The clinical care ran the other way. Hospitals with higher structural quality scores had better prophylaxis adherence — 95.5% against 93.3% — and worse measured VTE rates, 6.4 against 4.8 per 1,000.
The figures come from 2009-2010 Medicare claims for 954,926 surgical discharges at 2,786 hospitals, merged with 2010 Hospital Compare and American Hospital Association data on 2,838 hospitals. Bilimoria and colleagues published them in JAMA in 2013, and put the mechanism in one line: “Because they look more, they find more VTE events, paradoxically worsening their hospital's VTE quality measure performance.” The hospitals doing the better job were penalised by their own thoroughness. The instrument recording their outcome was their own decision to order the scan.
Look harder and the recorded rate nearly triples — 5.0 events per 1,000 in the lowest imaging quartile, 13.5 in the highest — while the prophylaxis runs the other way.
Steps
Trace every variable back to the decision that created it
The audit that prevents all of this is not a statistical procedure, and it happens before any model is fitted. Take each variable that carries weight — the treatment, the outcome, every confounder you intend to adjust for — and follow it backwards from the column in the table to the moment a person or a machine produced that value.
Write the construct in one sentence, in the words you would use talking to a colleague. Write the recording rule in one sentence, in the words the data dictionary uses. Then read the two sentences next to each other.
Most of the time they differ. The job of the audit is to make you say how they differ, in which direction, and for whom.
This is no longer a private discipline. A regulator now asks for it in writing. FDA's final guidance on using electronic health records and medical claims data to support regulatory decisions was issued on 25 July 2024, finalising a draft from 2021. It tells sponsors that how much validation a study variable needs depends on the consequences of misclassifying it. Listing what had changed between draft and final, the notice includes “recommending the use of quantitative approaches, such as quantitative bias analyses, either a priori for feasibility assessment, or to facilitate interpretation of study results, or for both purposes, to demonstrate whether and how misclassification, if present, might impact study findings”. The audit, and the sensitivity analysis that follows it, are what a regulator now expects to be shown. Not a scruple the analyst may or may not have.
- 1
Name the construct
Treatment receipt, disease state, preference, severity, or outcome.
- 2
Describe the instrument
Code, sensor, survey, observer, or administrative process.
- 3
Identify error drivers
Treatment, outcome, site, time, incentives, and missingness.
- 4
Find validation evidence
Gold standard, adjudication, repeat, or external source.
- 5
Analyze sensitivity
Bias correction, probabilistic measurement, or bounds.
Key idea
Adjusting for a noisy proxy can create false reassurance
A proxy that balances beautifully across arms is the most persuasive wrong thing in an analysis. Balance is a claim about the recorded column, and the recorded column is not the construct. The variation the proxy never captured can be badly unbalanced while the table shows nothing at all, because the check is measuring the part you already knew about.
A commercial risk-prediction algorithm applied to millions of US patients was meant to identify who needed extra help. What it predicted was health-care cost. Cost is a defensible stand-in for illness and it performs well by ordinary accuracy checks. It also encodes who gets care, not who is sick. Black patients at a given risk score were considerably sicker than White patients at the same score, and correcting the target would raise the share of Black patients flagged for extra help from 17.7% to 46.5%. The dissection ran in Science in 2019, and Obermeyer and colleagues drew the general lesson: “We suggest that the choice of convenient, seemingly effective proxies for ground truth can be an important source of algorithmic bias in many contexts.” Nearly two thirds of the eligible population was sitting in the part of the construct the proxy never touched. No diagnostic run on the recorded column would have said so.
The instinct at that point is to add more proxies. Sometimes it helps. It can also drag in variables the treatment itself affected, open a path through a collider, and widen the amount of personal data being held for no gain the analysis can demonstrate.
Three things do help, and none of them are free. Write down the construct model — what you believe the true variable is, and how the recorded one gets produced from it. Compare proxies against validation data wherever any exists. Then run a quantitative sensitivity analysis asking how much residual error it would take to move the conclusion, and print the answer beside the estimate rather than in an appendix.
A balance table can only report on the column you handed it.
Visual
The lineage from construct to recorded field
The diagram lays the chain out in the order it actually happens, which is not the order the analyst meets it.
It starts by defining the construct, the scientific quantity the study is about. Then it specifies the observation process: who records, when, and prompted by what. Then it maps the error pathways as arrows — which variables influence the mistakes, and whether any of those variables are the treatment or the outcome. Then validation, where the recorded value is set against a better measurement on some subset. Then propagation, where the uncertainty accumulated across all of that is carried into the final interval instead of being dropped at the moment the tables are merged.
The fourth stage is the one that can be watched working. A nutrition study called OPEN — Observing Protein and Energy Nutrition, run by the National Cancer Institute — measured 484 healthy volunteers from Montgomery County, Maryland, between September 1999 and March 2000. Energy intake came from doubly labelled water, protein intake from urinary nitrogen. Neither reference depends on anyone remembering what they ate. Alongside them the study ran a food frequency questionnaire and 24-hour recalls, the instruments nutritional epidemiology normally lives on. Kipnis and colleagues reported the gap in 2003: “Accounting for the reference biomarkers, the data suggest that the FFQ leads to severe attenuation in estimated disease relative risks for absolute protein or energy intake (a true relative risk of 2 would appear as 1.1 or smaller).” A doubling of risk arriving at the analyst as a tenth of one.
The study's second finding is the one that decides whether a validation stage is worth anything. Using the 24-hour recall rather than the biomarker as the reference underestimated the true attenuation by up to 60%. A validation reference that shares the questionnaire's failure mode reports back that nothing much is wrong.
Studies rarely break at the last stage. They break at the second, where the observation process was never written down — which means the arrows at the third stage cannot be drawn, and the fourth stage has nothing to check against.
- 1
Define construct
What scientific or operational quantity matters?
- 2
Specify observation process
Who records it, when, and with which instrument?
- 3
Map error pathways
Dependence on treatment, outcome, site, or selection.
- 4
Validate
Reference sample, repeated measure, calibration, or adjudication.
- 5
Propagate uncertainty
Correction models, sensitivity, and scope limits.
When measurement uncertainty changes the causal claim
Three situations change what you are entitled to claim.
If the recorded variable corresponds to assignment rather than receipt — the prescription order, not the swallowed pill — then say so in the claim itself. That instruction has a formal vocabulary, a date and a regulator behind it. ICH E9(R1), the addendum on estimands and sensitivity analysis in clinical trials, was adopted on 20 November 2019, and FDA announced its adoption as guidance for industry on 12 May 2021. Its glossary defines an estimand as “A precise description of the treatment effect reflecting the clinical question posed by the trial objective.” It defines intercurrent events. It sets out the treatment policy strategy, in which the value of the outcome variable is used regardless of whether the patient continued the treatment. Under that vocabulary the effect of being prescribed something is not a compromise forced by the data. It is a named estimand, chosen in advance. It becomes misleading only when it is given the other one's name.
If the outcome is detected differently across arms, no adjustment repairs it afterwards, because the missing events were never recorded to begin with. The hospitals in the lowest imaging quartile did not have fewer clots; they had fewer scans. That is a design fix: blinded assessment, or the same measurement schedule for everyone, decided before collection rather than argued over after.
And where validation is weak or absent, resist presenting the cleaner construct anyway. Report the effect for the operational definition you actually used. Then show how the conclusion moves across the range of measurement error a reader could plausibly believe — the quantitative bias analysis FDA asked for in the 25 July 2024 guidance. If the finding survives that range, you have something. If it does not, you have learned the most important fact available about your data: your answer is currently a property of the recording rule.
Name the variable you really recorded and a reader can judge the gap for themselves; name the one you hoped for and they cannot.
Key takeaways
- Every causal variable in a study is a recording rule standing in for a scientific quantity, and the two are never quite the same thing.
- Measurement error does not only shrink effects. Correcting single baseline blood pressure readings for regression dilution across 420,000 individuals made the associations about 60% greater; other error structures inflate, reverse or manufacture effects instead.
- Error that depends on the treatment or the outcome is the dangerous kind: hospitals in the highest VTE imaging quartile recorded 13.5 events per 1,000 against 5.0 in the lowest, while their prophylaxis adherence ran the other way.
- Adjusting for a proxy removes only what the proxy captured — an algorithm that predicted cost instead of illness flagged 17.7% of Black patients for extra help where correcting the target gave 46.5%.
- Validation samples turn a suspicion about error into a quantity: OPEN's doubly labelled water and urinary nitrogen showed a true relative risk of 2 arriving through a food frequency questionnaire as 1.1 or smaller.
- Measurement uncertainty belongs in the final causal conclusion — FDA's 25 July 2024 real-world data guidance recommends quantitative bias analysis to show whether misclassification would move the finding.