Research
When benchmark inferences do not compose: Projectibility in AI evaluation
Overview Research area: AI evaluation and measurement theory — specifically benchmark validity, generalization, and the logic of combining evidence across studies. Technical level: Advanced. The paper
- arXiv
- 2607.26159
- Published
- 2026-07-28
- Authors
- Brett Reynolds
AI summary
Overview
Research area: AI evaluation and measurement theory — specifically benchmark validity, generalization, and the logic of combining evidence across studies.
Technical level: Advanced. The paper builds on argument-based validity theory (Kane, Messick), Goodman's problem of induction, generalizability theory, and causal transportability, and it uses typed notation and a probability identity. It is written for measurement researchers and evaluation methodologists rather than for general readers.
Scope (one sentence): The paper argues that separately warranted benchmark inferences do not automatically form a warranted chain, and it supplies a "projectibility audit" — a link-by-link procedure for diagnosing where benchmark-to-use arguments fail to join.
What This Paper Is About
An AI benchmark result is almost never used on its own. Someone generalizes it to further cases, reads it as evidence of a capability, extrapolates it to new tasks, moves it to another system or site, and combines it with assumptions about human review and downstream consequences. Existing validity-centred approaches require evidence for each of those claims, but they leave a prior question to the analyst: whether the end of one study is actually the beginning of the next. This paper makes that interface problem explicit, gives it a formal statement (a "non-composition principle"), and provides a procedure for auditing it.
Key Contributions
-
A non-composition principle. The paper states that warrant for claim C₁ and warrant for claim C₂, each established separately, do not yield warrant for the composed inference C₂ ∘ C₁. This is presented as a principle about inference, not a theory of warrant.
-
A typed representation of empirical endpoints. Each empirical "node" is written as a record with five fields — object, population, conditions, outcome, and period — and each projection is an "edge" running from a source node to a target node with its own projective claim, assumptions, and evidence. This forces the proposed match between two studies to become inspectable instead of being asserted through a shared broad noun.
-
A seven-part interface audit. Five endpoint-alignment requirements (object continuity, population alignment, condition alignment, outcome and scale continuity, temporal alignment) test whether the links meet at all; two further requirements (compatibility of assumptions and effect modifiers, propagation of dependence and uncertainty) test whether warrant transmits once they do. The procedure separates links that never meet from links that meet while warrant fails to cross.
-
Three-way distinction among composition, convergence, and replacement, with the argument that "benchmark plus local validation" describes all three and that only one of them is a composable chain.
Main Findings
-
Composition is a distinct problem from link validity. Validity-centred approaches assess whether evidence supports a stated benchmark interpretation and use. The paper's claim is that this leaves open whether one supported conclusion supplies the premise the next link requires. The paper states directly that the non-composition principle is not a gap in argument-based validity — Kane's framework already has room for interface assumptions — and that the audit's contribution is making their identification systematic.
-
Shared labels supply no warrant. A benchmark containing "legal tasks" and a firm performing "legal tasks," or a developer testing "the model" and a deployer using "the model," conceal the gap. Repetition of a broad label does not establish that the tested and deployed objects are continuous.
-
Spurious adjacency counterexample. The paper presents two individually true claims: a model answers 90% of held-out supplied-text legal questions correctly under the benchmark protocol, and lawyers following a written review instruction catch 99% of fabricated citations in drafts produced by a retrieval-based application for routine research requests. Neither is defective alone, but their conjunction does not warrant that reviewed memos are substantively correct. If the application omits a controlling authority on 20% of requests and the lawyers' procedure requires no independent update search, one fifth of final memos contain a substantive defect while both premises remain true. The chain fails at two interfaces: the outcome changes from supplied-text correctness to citation fabrication, and the review evidence does not cover omitted authority.
-
Positive evidence would look different. The firm could prepare an authority list independently of the system, score omitted controlling authority as well as citation fabrication, and require reviewers to compare every final memo against that list or conduct a defined update search. This would supply observations for the missing workflow link without making the benchmark general.
-
A real-world instance. Legal-research vendors advertised retrieval-grounded products as delivering "100% hallucination-free linked legal citations" or as avoiding hallucination by relying on trusted content, and Magesh et al. (2025a) found that no empirical evidence accompanied those claims. Their preregistered evaluation reports hallucination on more than 17% of queries for both the LexisNexis and Thomson Reuters tools, and incomplete answers on more than 60% for Thomson Reuters. The paper notes that incompleteness is the omission failure, not the fabrication failure the marketing addressed.
-
Transmission failure is harder to see because the links do meet. The paper's second counterexample has several systems evaluated on the same held-out requests, a benchmark score predicting the frequency of draft defects across them, and reviewers working from one instruction catching 90% of draft defects across drafts from all of those systems. The presented text is truncated at this point.
-
What composition requires quantitatively. With B a benchmark result, D a draft-level outcome, and Z a final outcome after review, the paper writes P_T(Z | B) = Σ_d P_T(Z | D = d, B) P_T(D = d | B). One study may estimate P_T(D | B) and another P_T(Z | D), but the composition requires P_T(Z | D, B). Substituting the second quantity is warranted only if Z ⟂ B | D — the benchmark result adds no information about review success once the draft outcome is known. Otherwise the substitution rests on discrepancies happening to cancel in the weighted sum, which the paper says an evaluator is usually not in a position to claim. The independence requirement could fail if, for example, high-scoring systems produce more fluent errors that reviewers are likelier to trust.
-
Dependence must be propagated, not reset. The paper notes that twenty outputs from one build are not twenty systems, and thirty-two comparisons among related models and shared benchmarks are not thirty-two independent replications.
-
Interpretive claims are not empirical endpoints. A capability attribution such as "general reasoning" is an interpretive claim about a bearer, not a sampled endpoint. It may serve as an interpretive premise for a projection, but it does not carry a sampled population of cases, a criterion outcome, or the conditions under which the attributed capability is expected to predict that outcome.
-
Measured convergence is not downstream behaviour. Jung et al. (2026a) applied human psychometric instruments to language models and reproduced the inter-test correlations the theory predicts: sexism with racism at r_s = .47, authority with benevolent sexism at .43, and fairness with hostile sexism at −.37. The paper says those agreeing measurements bear on what the instruments measure, but not on how the models then behaved in matched downstream settings.
-
Adjacent frameworks stop at different points. Raji et al. (2021a) argue that broad benchmark claims outrun the contextual tasks used to construct them; Bowman & Dahl (2021a) set out how construct validity fails in language-model evaluation; Liu et al. (2024a) adapt evidence-centred design to benchmark construction; Salaudeen et al. (2025a) offer a claim-aware framework mapping measurements to claims; Freiesleben & Zezulka (2025a) specify measurement conditions for scientific inference; Bean et al. (2025a) document construct-validity weaknesses across 445 language-model benchmarks. None, on the paper's account, guarantees that the target of one analysis is the source of the next.
Methodology in Plain English
The paper is theoretical and conceptual rather than experimental. It proceeds by:
-
Locating the problem in existing validity theory. The paper adopts Kane's view that validity belongs to a proposed interpretation and use of scores, and Messick's unified account, then identifies the composition question those frameworks leave to the analyst.
-
Defining terms precisely. A projection extends an interpretation, prediction, or explanation from specified source observations to unobserved cases; projectibility concerns whether that bounded extension is warranted relative to a declared target and use.
-
Building a notation. Each empirical study endpoint is represented as a five-field record (object, population, conditions, outcome, period) and each projection as an edge carrying a claim, assumptions, and evidence. This makes "do these two studies meet?" a checkable comparison of fields rather than a narrative assertion.
-
Deriving the audit requirements. Five alignment requirements come from the five fields; two transmission requirements come from the assumptions and evidence the edges carry. The paper says both sets presuppose empirical nodes at both endpoints.
-
Testing the procedure against worked cases. Two counterexamples — one where the links never meet (spurious adjacency) and one where they meet but warrant does not transmit — show the two failure modes coming apart under inspection. The paper describes a legal-research case in which hypothetical results support one projection, defeat another, and leave a third unresolved, and a known-truth demonstration showing why aggregate stability can erase distinctions a later projection requires.
-
Comparing to neighbouring frameworks. A table sorts argument-based validity, nomological networks, estimand specification, generalizability theory, and causal transportability by their primary object, the question each answers, and the further composition question each leaves open.
Why This Matters
Impact on research. The paper argues that the failure mode is not a rare corner case. AI makes the interface problem pressing because one evaluated component can be replicated at high volume, modified into many applications, and inserted into workflows whose relevant evidence is distributed among different actors; an unsupported bridge can become a repeated, correlated failure while no single study observes the complete chain. The paper also states two obligations that follow and fall outside any single link: a source report has to preserve the resolution a declared downstream claim will need, because an aggregate can leave several target-relevant states indistinguishable; and the evidence divides between developers and deployers, because no single party observes the whole chain.
Real-world applications:
- Legal research tools. The Magesh et al. (2025a) case shows how an inference from a component property (retrieval over an authoritative database) to what the product asserts went unaudited, with hallucination on more than 17% of queries and incomplete answers on more than 60% for one vendor.
- Enterprise AI procurement. A firm deciding whether to adopt a benchmark-topping model for professional work can use the audit to check whether its own workflow evidence composes with the vendor's benchmark evidence or instead replaces it.
- Human-in-the-loop review design. The transmission requirement forces explicit attention to whether review success depends on the draft outcome alone, or also on system quality — the Z ⟂ B | D condition.
- Regulatory and policy arguments about AI capability, where a chain from benchmark scores to general reasoning to deployment consequences is often presented as if the links composed.
Industry relevance. The paper's distinction between composition, convergence, and replacement has direct operational consequences: a local study evaluating only one frozen application has a constant benchmark profile across its requests, so it replaces rather than confirms the benchmark-to-use projection. Evaluating several systems and testing whether their benchmark score profiles predict defect rates on the same requests is the version that tests a composable predictive edge.
Future Directions
- Bridging studies. The audit identifies which field mismatch leaves a bridge owed; a next step is designing those bridges, whether as a widened generalizability study spanning both declared universes or as direct target sampling.
- Where causal identification is required. The paper states that projectibility claims no identification theorem, and that where the target claim is causal a projectibility audit should defer to the corresponding causal requirements rather than treating generic robustness evidence as enough.
- Resolution-preserving reporting. The paper identifies an obligation for source reports to preserve the resolution a downstream claim will need, since an aggregate can leave target-relevant states indistinguishable. What form that reporting should take, and how it interacts with the known-truth demonstration in Section 5, is left open in the provided text.
- Dividing evidential work between developers and deployers. Sections 6 and 7 are described as applying the result to general-capability claims and dividing the evidential work between developers and deployers, but the full text of those sections is not included in the content provided.
Target Audience
Evaluation methodologists and measurement researchers working on AI benchmark validity; AI evaluation practitioners who assemble evidence from multiple studies or vendors into a single capability or deployment claim; and researchers in psychometrics, educational measurement, and causal inference interested in how their frameworks do and do not compose across separately conducted studies. Readers who need only practical benchmark guidance rather than the underlying inferential architecture would likely be served better by the claim-centred frameworks (Salaudeen et al., Freiesleben & Zezulka, Bean et al.) that this paper builds on.
Note on scope of this summary: the provided paper content is truncated during Section 3.4, so details of Sections 4 (the legal-research case), 5 (the known-truth demonstration), 6 (general-capability claims), and 7 (developer/deployer division) are reported here only as the paper's own introduction and abstract describe them.
Authors’ abstract
An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper's distinctive claim is a non-composition principle: support for adjacent projections warrants their composition only when endpoints and assumptions align and dependence and uncertainty are carried through. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A reanalysis and simulation show why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.