Skip to content
AI.info

Research

What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals

Overview Research area: Software engineering for AI evaluation infrastructure (cs.SE), sitting at the intersection of benchmark reproducibility, measurement validity, and partial identification. Techn

What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals
arXiv
2608.19269
Published
2026-08-18
Authors
Xi Qin, Jizhou Tong

AI summary

Overview

Research area: Software engineering for AI evaluation infrastructure (cs.SE), sitting at the intersection of benchmark reproducibility, measurement validity, and partial identification.

Technical level: Intermediate. The formal core (identified sets over evaluator families) is compact, but the paper's real weight is in its audit protocol, its census accounting, and its distinction between evidence availability, evaluator meaning, and claim resolution. Readers need comfort with rankings, metrics, and the idea of a specification curve; no deep statistical theory is required.

Scope in one sentence: The paper freezes a large community evaluation collection (Inspect Evals) at one repository commit and asks, for each stated claim, which answers remain true across the evaluator meanings admitted to the analysis — and when the evidence needed to answer at all is simply absent.

What This Paper Is About

Benchmarks can run without telling you what their published results actually license. The authors take a frozen snapshot of Inspect Evals, a community-contributed evaluation collection and register for the Inspect AI framework, and try to replay its historical claims from the evidence preserved in the repository.

The core problem is that inference is left "ambient": it depends on repository layout, maintainer knowledge, README conventions, implicit missingness rules, and manual evidence assembly. The paper makes that inference step explicit and executable, so that a claim either returns supported answers across admitted evaluator readings — or stops with a named reason and the minimum artifact that would unblock it.

Key Contributions

  1. A claim-replay formulation. The authors define a contract over frozen evidence D, an admitted evaluator family F, and a stated claim q, and compute a claim-relative identified set — the collection of every answer q(V(f;D)) allowed by the admitted specifications. The claim is identified when that set has one member.

  2. A complete commit-bound census. At public commit 32b79a2, a predeclared mechanical eligibility predicate yields 124 units from 129 strict-path candidates (src/inspect_evals/<id>/eval.yaml), with five enumerated exclusions. Fourteen units permit deterministic claim analysis; 110 stop with an explicit reason.

  3. A practical audit and release interface. The interface reports exact values, directions or threshold crossings, winners/top sets, complete weak orders, and stable pairwise relations separately, each accompanied by a disagreement witness when answers differ, and a typed stop with the first missing item otherwise.

  4. A separation of what are usually conflated. The paper distinguishes evidence stopping (an executability state), semantic multiplicity (a family property), and claim-resolution separation (an inference result), and keeps evaluator variation, sampling uncertainty, printed precision, and reversible unit changes as distinct coordinates.

Main Findings

  • Most units cannot be replayed at all. Of the 124 eligible units, 110 stop and 14 permit deterministic analysis. Of the 110 stops, 103 belong to the outcome-blind census-review path (114 cases) and seven to the frozen random draw.

  • The leading stopping reasons are binding failures, not bad benchmarks. On the outcome-blind path, the primary reasons are evidence-binding failure (47), missing comparative observations (35), and missing judge traces (9); the remaining twelve concern endpoint or target definition, missingness, claim selection, or terminal state. Each record names the first unavailable item and the minimum artifact that would permit claim replay.

  • The frozen random draw completes at 3/10. The draw allocates two units per stratum from stratum sizes (47, 40, 26, 7, 4). The paper states this is the observed completion fraction in that frozen balanced-strata draw, not a unit-weighted frame rate and not a rate of non-identification.

  • One result, two documented readings. On the same AgentDojo attack outcomes, the published-micro reading ranks Gemini 2.0 ahead of GPT-4o-mini (20.827 vs. 27.186; lower is better), while the utility-conditional reading reverses them (32.000 vs. 19.745). Claude 3.5 remains first under both. The score and the lower ordering change; the winner does not.

  • Claim level determines what survives. In the AgentDojo primary analysis, there is one winner, one full order, and 3/3 pairs agree. In the wider review family, Claude still wins, but there are two full orders and 2/3 pairs agree. AutoML shows the same pattern: the primary family is identified with 10/10 stable pairs, while the review family is not identified with 9/10.

  • Expanding the family need not erase everything at once. Across the three purposive audits (AgentDojo, AutoML, BEIR), 13 of 15 previously identified pairwise relations remain identified when moving from the primary to the review family, and no previously identified winner is lost.

  • The completed random draws are non-identified. All three completed random draws yield more than one winner and more than one full order in both families, with stable-pair counts of 0/1, 0/1, and 0/3 in both. Their mechanisms differ: credit definition, aggregation, and claim target.

  • A separate BEIR diagnostic shows a larger collapse. The 25-view coverage analysis is not part of the named-family result. It is an exploratory coverage sensitivity in which the winner varies and stable pairs fall from 2/3 to 0/3. The paper explicitly does not license the enlarged grid as a completed or admitted family.

  • Changing one terminal rule collapses an effect. In a controlled 217-row diagnostic dataset outside the census denominator, rows, support, and seven prediction rules are fixed and only the terminal rule changes. Replacing the original one-sided rule with either symmetric rule moves the adoption effect from approximately -40 percentage points to zero. Exact original effects span -43.03 to -40.76 pp across the three disposition policies.

  • A fixed winner can coexist with a varying order. Under the exclude-both specification, the adoption effect is exactly zero, the complete ranking spans five weak orders, and Mini is the unique winner in every cell. The paper treats effect, full order, and winner as different questions about the same observations.

  • Disagreement is not attributable to one challenger. The observed disagreement persists under 71 of 76 leave-one-out removals, and admitted–admitted pairs alone furnish 14 distinct disagreement witnesses.

  • The pipeline fails closed. In a controlled evidence-ablation test over five otherwise complete packets, all five intact packets remained complete, while all 20 targeted evidence deletions stopped exactly at the declared G1–G4 boundary. The verifier includes an independent 844-check path.

Methodology in Plain English

The authors fix a snapshot of an evaluation repository at a single commit so the audited frame is reproducible, then apply a mechanical eligibility rule to decide which evaluation units are in scope. Nothing in that rule depends on how any benchmark actually performed.

For each eligible unit they try to reconstruct a historical claim. Reconstruction requires evidence: sample-keyed outputs, judge decisions or sufficient terminal state, support and missingness decisions, and a declared evaluator meaning. If any of these is unavailable, the unit receives a typed stop record naming the first missing item rather than a substituted number or a fresh run. A fresh model, judge, agent, or sandbox run changes the evidence D and cannot reconstruct the historical event, so the stops stay in the census.

For units where the evidence is sufficient, the authors enumerate the evaluator specifications admitted to a family. Reviewers classify each candidate as admitted for the primary analysis (A), disputed but relevant for review (D), or excluded (X); the primary family contains A specifications and the review family contains A plus D. Each specification is expanded into a semantic map, a row rule, and an aggregation rule, which forces eligibility, row scoring, and aggregation to be declared separately. Then they compute the answer under every admitted specification and compare.

The comparison is done claim by claim. A claim can ask for an exact value, a direction or threshold crossing, a winner or top set, or a complete weak order. Because adding evaluator meanings can only enlarge the answer set, a conclusion fixed in the primary family may vary in the wider review family, and a coarse claim can be fixed while a finer one varies. Pairwise relations are collected separately, and a pair is "stable" when every admitted specification agrees on it. Where answers differ, two specifications with different answers serve as a concrete witness.

Three evidence streams are kept separate and never pooled: the complete 124-unit census, a frozen random draw of ten selected before terminal audit outcomes, and a mechanism portfolio of three purposive deep audits plus eight labeled outcome-exposed discoveries. The paper also notes that generative AI tools, principally OpenAI Codex, assisted with refining the framework, checking mathematical claims, implementing and testing code, cleaning bound data, and drafting, while no AI system had final authority over claim selection, semantic admission, evidence licensing, interpretation, or reporting.

Why This Matters

Impact on research. The paper argues that benchmark releases should ship a claim-replay interface alongside task code and a headline metric, binding observations or a sufficient statistic, evaluator or judge version, support and missingness decisions, aggregation, the admitted evaluator family, and the claim being asked. It also argues that a flip-only or single robustness label discards decision-relevant structure: it hides stable winners and stable pairwise relations while overstating the generality of what varies.

Real-world applications:

  • Leaderboard and model-card reporting. A release could report the identified set for its primary family and its wider review family side by side, with witnesses and stable comparisons, so readers can see which rankings are contingent and which are not.
  • Benchmark maintenance and archival. The stopping taxonomy gives maintainers an actionable list: for each unit, the first unavailable artifact and the minimum addition that would make a historical claim replayable.
  • Audit and procurement review. An external auditor can inspect the semantic ledger — source locators, quotations, same-target and support rationale, A/D/X judgments, and the evidence horizon — even though the automated verifier does not prove construct validity.
  • Prospective evaluation design. Teams registering new evaluations before outcomes could freeze the candidate family and claim set in advance, then measure replay completion against the legacy interface.

Industry relevance. Evaluation frameworks, registries, structured logs, and containers are now standard infrastructure, but the paper's point is that running code is not the same as replaying a claim. The audit at one commit found that task and scorer availability repeatedly precedes missing historical observations, judge decisions, state, support joins, or an evaluator registry — meaning the audited artifacts often support forward execution without carrying the last links needed to reconstruct a published conclusion. That is a supply-chain-style discipline argument applied to evaluation conclusions.

Future Directions

  • A prospective utility test. The retrospective census does not estimate how many of the 110 stops a new release process would prevent. The paper proposes registering new units and advertised claims before outcomes, requiring claim-replay fields at release time, and comparing replay completion, time-to-audit, and decision changes against the legacy interface.

  • Claim-relative preservation standards. The paper's preservation chain shows that a retained scorer is only the first layer; later claims may require historical outputs, mediator decisions, state, support joins, and an admitted semantic registry. How to specify the least artifact that closes a declared query — a compact sufficient statistic for a direction versus sample-keyed outputs for an exact value — remains open.

  • Scaling beyond one frozen frame. The 124-unit frame is the complete eligible frame at that commit, not the full distributed Inspect ecosystem, and the results are explicitly local to the named finite frame, evidence stream, frozen evidence, and enumerated evaluator family. Whether the same accounting holds across the distributed register is untested.

  • Whether a claim ledger changes practice. The paper does not claim causal utility, ecosystem prevalence, independent replication, benchmark validity, or future stop reduction. Whether exposing identified sets actually changes what reviewers and readers conclude is an open empirical question.

Target Audience

This paper is most useful to benchmark maintainers and evaluation infrastructure engineers who publish results and want to know what evidence their releases actually preserve; to measurement and reproducibility researchers working on specification curves, multiverse analyses, and partial identification in machine learning evaluation; and to audit, policy, or procurement reviewers who need to judge whether a published model comparison survives reasonable variation in evaluator meaning. Readers interested in the formal side will find the object compact, but the paper rewards readers who care about release engineering, evidence binding, and failure reporting at least as much as about statistics.

Authors’ abstract

Benchmarks can run without determining what their results license. We freeze a large evaluation collection and attempt to replay its historical claims. Most units stop because the evidence required for replay is not bound. Where replay is possible, different claims remain stable at different resolutions. We make this otherwise implicit inference step explicit and executable.

Read the original paper