Research
Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation
Overview Research area: Benchmark auditing and evaluation methodology for agentic AI systems — specifically contamination susceptibility in public coding benchmarks and leaderboards. Technical level:

- arXiv
- 2610.05830
- Published
- 2026-10-05
- Authors
- Dishu Yang, Qi Su, Hongbo Qin, Hansong Zhang
AI summary
Overview
- Research area: Benchmark auditing and evaluation methodology for agentic AI systems — specifically contamination susceptibility in public coding benchmarks and leaderboards.
- Technical level: Advanced. The paper combines provenance auditing, matched-control causal design, cluster-aware uncertainty estimation, and scorer validation with formal estimand definitions.
- Scope: A "measurement-first" audit framework applied to nine Holistic Agent Leaderboard (HAL) configurations and to a reported file-localization gap on SWE-bench Verified, evaluated on GPT-4.1 and DeepSeek-V4-Flash.
What This Paper Is About
Public coding benchmarks such as SWE-bench Verified keep their task statements and solution-bearing artifacts publicly accessible, so a high leaderboard score may reflect recall of known solutions rather than genuine problem-solving. The authors argue that a single "contaminated or clean" verdict conflates claims that demand very different kinds of evidence, and they build a framework that ties the strength of a contamination claim to the evidence actually available for it. Their goal is not to prove contamination but to determine which claims the evidence can and cannot support.
Key Contributions
- A three-channel threat framework with a fail-closed coding rule. The framework separates training-time exposure, evaluation-time retrieval, and pipeline/scaffold leakage, each coded open (O), partial (P), closed (C), or unknown (?). Missing evidence is explicitly not treated as negative evidence, so no channel can be declared closed without complete negative evidence within scope.
- A structured susceptibility and incident card applied across nine HAL configurations. Each row records in-scope configurations, channel labels, linked evidence at a pinned revision, and confirmed incidents separately from susceptibility.
- An outcome-blind, same-repository matched-control behavioral design. SWE-bench Verified tasks are matched one-to-one without replacement to merged bug-fix pull requests from the same repository, matched on timing, issue size, and pull-request complexity, with symmetric prompt-leakage screening, repository-aware uncertainty analyses, and a post-hoc pair-integrity audit.
- Validation of the reproduction scorer as a gate on inference. Probe 2 contrasts only adjudicate hypotheses if the wrong-gold audit supports low false-positive behavior, correct-gold firings are confirmed by consensus human labels, and design-weighted review supports sufficient sensitivity. Inconclusive validation does not count as a pass.
Main Findings
- No channel coded closed across 27 assessments. Across the nine HAL configurations, none of the 27 channel assessments was coded closed, yet incidents were confirmed in four configurations.
- Channel labels were broad, incidents were narrow. Training-time exposure was partial for all nine configurations; evaluation-time retrieval was open for seven and partial for two; pipeline/scaffold leakage was open for four, partial for two, and unknown for three. Confirmed pipeline incidents occurred in AssistantBench, SciCode, and TAU-bench Airline (all current revision) and in USACO (historical). No configuration-scoped retrieval incident was established.
- All training-time partial labels reflect the aggregation rule, not homogeneous evidence. Temporal eligibility, which is model-scoped, was unknown for every configuration; eight of the nine had affirmative support for the other three prerequisites, and Online Mind2Web additionally lacked public gold and task-gold linkability.
- ScienceAgentBench received a bounded "none" judgment. Online Mind2Web, GAIA, CORE-Bench Hard, and SWE-bench Verified Mini remained unknown. The HAL repair record documents 29 GAIA mirror visits for Generalist runs, but the audited GAIA row covers Open Deep Research, so the judgment was left unknown rather than transferred across scaffolds.
- Discoverability demonstration. In a focused audit of 20 SWE-bench Verified Mini tasks with 40 frozen queries against one web-search service on August 25, 2026, 35 of 40 queries returned a task-linked result in the top ten. Gold information was directly reachable for 7 queries, partially reachable for 23, and not reachable for 10. The authors state these counts are not a prevalence estimate.
- Probe 1 gaps were positive but inconclusive. Among the 100 pairs retained after symmetric screening and the pair-integrity exclusion, GPT-4.1 showed a +10.0-point pair-weighted Top-3 benchmark-associated gap, but the 95% intervals from both the prespecified paired-bootstrap procedure and the post-hoc repository-balanced analysis included zero.
- Probe 1 primary accuracy figures. Top-3 accuracy was 0.640 versus 0.570 for DeepSeek-V4-Flash and 0.710 versus 0.610 for GPT-4.1 (benchmark versus control). Pair-weighted gaps were +0.070 (95% CI [-0.070, +0.210]) and +0.100 ([-0.030, +0.230]); repository-balanced gaps were +0.081 ([-0.079, +0.240]) and +0.109 ([-0.116, +0.334]).
- Hierarchical intervals also included zero. Resampling repositories and then pairs within repositories gave Top-3 95% intervals of [-0.132, +0.283] for DeepSeek-V4-Flash and [-0.128, +0.334] for GPT-4.1.
- Secondary-metric asymmetry for GPT-4.1. Three secondary-metric intervals nominally excluded zero under the prespecified pair-weighted analysis, whereas none did under the post-hoc repository-balanced analysis. All DeepSeek-V4-Flash intervals included zero under both analyses.
- Directional consistency, not magnitude comparability. Within the analysis subset, GPT-4.1 and DeepSeek-V4-Flash showed Top-3 gaps of +10.0 and +7.0 points, directionally consistent with Liang et al.; the different cohorts and metrics preclude a magnitude comparison.
- Leakage rates were similar across arms. At least one gold source path or basename appeared verbatim in 63 of 199 benchmark issues (31.7%) and 57 of 199 control issues (28.6%). Excluding both members of any leaking pair removed 98 pairs (49.2%), leaving 101; the cross-arm underlying-issue exclusion left 100 pairs for Probe 1. The 31.7% benchmark-arm rate is close to the 27.0% reported by Liang et al.
- Pair-level exclusion varied widely by repository. Exclusion rates ranged from 26.7% (Sphinx) to 76.0% (scikit-learn), making unaudited results sensitive to repository composition. Because detection is verbatim, these rates are lower bounds.
- The leakage screen selects on issue characteristics. Retained benchmark issues were shorter (172 versus 328 words), more HAL-linked, less often task-type concordant, and differently distributed across repositories than excluded ones.
- The reproduction scorer did not pass its validation gate. Against consensus human labels, sufficient scorer sensitivity could not be established for either model. The strict scorer detected 0 of 2 consensus-positive DeepSeek-V4-Flash responses and 1 of 3 consensus-positive GPT-4.1 responses, corresponding to design-weighted sensitivity estimates of 0.00 and 0.19, with unweighted exact 95% intervals of [0.00, 0.84] and [0.01, 0.91].
- False positives on correct-gold comparisons. Both DeepSeek-V4-Flash firings on correct-gold comparisons were false positives. The single GPT-4.1 firing was human-positive. Precision was 0/2 for DeepSeek-V4-Flash and 1/1 for GPT-4.1.
- Scorer agreement statistics. After dichotomizing ratings at the >= 2 threshold, Cohen's kappa was 0.788 for DeepSeek-V4-Flash and 1.000 for GPT-4.1 within the enriched sample.
- Probe 2 was sparse. Only 42 of the 199 matched pairs were strictly scorable, and the scorer fired on just 3 of 168 model-arm outputs.
- Hypothesis disposition. H1 is inconclusive: all point estimates were positive, but no Top-3 primary interval excluded zero. H2 and H3 are unresolved because the scorer did not pass the validation gate.
- Robustness. Leave-one-repository-out re-estimation shifted the four reported repository-balanced gaps by at most 0.068. In the post-hoc 64-pair task-type-concordant analysis, all Top-3 intervals included zero.
- No claim of membership or inflation. The results do not establish training-data membership, contamination prevalence, or benchmark-induced score inflation.
Methodology in Plain English
The authors split the audit into two stages and treat measurement quality as a precondition for any conclusion.
Stage I — provenance and incident audit. They assembled a structured "card" for nine HAL configurations, pinning each judgment to a specific revision with a fixed snapshot. The primary evidence search was frozen on August 25, 2026, and revision-pinned checks continued through September 9. Two authors independently coded 117 field-level judgments and agreed on 90 (76.9%) before adjudication; final channel labels were then derived mechanically from the adjudicated fields. Automated procedures only flagged candidate evidence; confirmation required human review of the source, benchmark identity, threat channel, and revision. Three supplementary artifact audits were run: a local inventory of task artifacts, prompts, retrieval mechanisms and evaluator boundaries; a 40-query search-discoverability audit over 20 SWE-bench Verified Mini tasks; and a comparison of local benchmark-test and non-test artifacts.
Building the matched-control cohort. They started from a pinned revision of the 500-task SWE-bench Verified test split, retained all 50 HAL Mini tasks, and sampled 150 more with deterministic repository-stratified sampling using seed 13. Controls came from 5,172 merged, issue-linked pull requests passing a frozen lexical bug-fix screen, after excluding all pinned SWE-bench Full, Lite, and Verified instances. Both arms required issue statements of 30 to 1,200 words, and repository quotas were capped at 30 tasks. Feasible pairs shared a repository, had merge dates within 365 days, differed by at most one changed file and 200 issue words, and shared the same temporal category relative to the documented GPT-4.1 cutoff of June 1, 2024. Minimum-cost one-to-one matching ran separately within repositories with no control reuse, leaving 199 pairs (50 HAL-linked, 149 expanded). A post-hoc pair-integrity audit removed one same-issue collision from both probes without replacement. Balance on ten covariates satisfied |SMD| < 0.20 in the pooled, HAL-linked, and expanded cohorts.
Two probes. Probe 1 gives the repository name and a sanitized issue description and asks the model to predict up to three production-source files modified by the accepted fix, scored against the complete accepted pull request with Top-3 accuracy as the prespecified primary endpoint. Probe 2 gives the repository, the sanitized issue, and the path of one modified file, and asks for the key code change, scored by exact-line recovery.
Symmetric leakage screening. Probe 1 drops a pair if either issue reveals a complete gold path or basename. Probe 2 drops a pair if, after whitespace normalization, either issue contains a non-comment production-code line added by the accepted patch with at least five lexical tokens and 20 non-whitespace characters. Both arms are removed whenever either is revealing. Matches found only in doctest or REPL examples are recorded separately and do not trigger exclusion.
Models and inference. GPT-4.1 and DeepSeek-V4-Flash were run at temperature 0, one completion per task, with no tools, browsing, or retrieval.
Statistics. The prespecified pair-weighted estimand averages benchmark-minus-control differences across matched pairs, with a paired percentile bootstrap using 10,000 replicates that resamples pairs but does not explicitly model repository-level dependence. The post-hoc repository-balanced estimand weights represented repositories equally, using a t-interval over repository-specific gaps (R-1 degrees of freedom; R >= 3), together with hierarchical repository-and-pair bootstrap and leave-one-repository-out sensitivity analyses. Analyses are conducted separately by model and metric, intervals are pointwise and unadjusted, and no family-wise confirmatory claim is made.
Scorer validation. Strict-scorable Probe 2 responses were rescored against deterministically selected, disjoint wrong-gold signatures from other tasks in the same repository. Two authors then reviewed a stratified enriched sample — all strict exact hits, all fuzzy-only hits, and the 15 highest-overlap outputs — blinded to benchmark-control membership but not to model identity or scorer-derived stratum, plus 20 outputs sampled without replacement from the remaining tail with seed 2026 and weighted by inverse inclusion probability. Consensus labels of 2 or 3 on a 0-to-3 scale indicated patch-specific implementation evidence.
Why This Matters
- Impact on research: The paper reframes contamination auditing from a binary verdict to an evidence-graded claim, and it demonstrates that a widely used behavioral probe cannot support its intended inference when its scorer is not validated. It also shows that a leakage screen can itself reshape the sample, limiting inference to a shorter-issue subpopulation.
- Real-world applications:
- Benchmark and leaderboard maintainers deciding what a score can legitimately be claimed to demonstrate.
- Agent developers and evaluation engineers who need to document revision-pinned provenance and retrieval conditions for their scaffolds.
- Procurement and model-selection teams that rely on public coding benchmarks as evidence of bug-repair capability.
- Research groups building new behavioral contamination probes who need a template for symmetric screening, matched controls, and scorer validation.
- Industry relevance: Organizations that publish or consume leaderboard results inherit the measurement problems described here. The paper's practical message is that leaderboard rankings from benchmarks with publicly accessible task statements and solution artifacts should be accompanied by provenance evidence, matched public controls, and validated scoring instruments before being treated as evidence of capability.
Future Directions
- Evaluate semantically aware reproduction scorers on larger, independently annotated samples, since the strict exact-match scorer missed most human-positive responses.
- Extend the matched-control design across additional repositories and model families, as the current study covers ten mature Python repositories and two API-served, non-reasoning models.
- Recover run traces and configuration attribution for the HAL configurations whose channel labels remain unknown, particularly Online Mind2Web, GAIA, CORE-Bench Hard, and SWE-bench Verified Mini.
- Address the design limitations the authors identify: controls are public tasks that precede the documented GPT-4.1 cutoff rather than confirmed unexposed negatives; no documented cutoff was available for DeepSeek-V4-Flash; manual review was blinded to arm but not to model identity or scorer stratum; and generation used no fixed API seed, so committed completions cannot be regenerated exactly.
- Resolve the open question of whether issue clarity is a plausible alternative explanation for the positive but inconclusive file-localization gaps.
Target Audience
Benchmark and leaderboard maintainers, agentic evaluation researchers, and contamination-auditing methodologists will benefit most. The paper is also relevant to practitioners who commission or interpret coding-benchmark evaluations, and to statisticians working on clustered, matched-control designs in machine learning evaluation. Readers need comfort with bootstrap and clustered uncertainty analysis, matched-cohort construction, and the distinction between susceptibility and realized incidents; newcomers will find the three-channel framework and the fail-closed coding rule conceptually accessible even if the statistics are not.
Authors’ abstract
Agentic leaderboards increasingly evaluate systems on public benchmarks whose task statements and solution-bearing artifacts can remain accessible. We propose a measurement-first audit framework that distinguishes contamination claims according to the evidence required to support them. It separates three channels that require different evidence: training-time exposure, evaluation-time retrieval, and pipeline/scaffold leakage. Each channel is coded as open, partial, closed, or unknown under a fail-closed rule. Across nine Holistic Agent Leaderboard (HAL) configurations, none of the 27 channel assessments was coded closed, but incidents were confirmed in four configurations. We then apply the behavioral component of the framework to a reported file-localization gap on SWE-bench Verified, using an outcome-blind, same-repository matched-control design with symmetric prompt-leakage screening, paired and repository-aware uncertainty analyses, and scorer validation, evaluated on GPT-4.1 and DeepSeek-V4-Flash. Among the 100 pairs retained after symmetric screening and the pair-integrity exclusion, GPT-4.1 showed a $+10.0$-point pair-weighted Top-3 benchmark-associated gap, but the 95\% intervals from both the prespecified paired-bootstrap procedure and the post-hoc repository-balanced analysis included zero, leaving the benchmark-associated gap inconclusive. The reproduction scorer did not pass its validation gate: against consensus human labels, sufficient scorer sensitivity could not be established for either model, and both DeepSeek-V4-Flash firings on correct-gold comparisons were false positives. Without provenance evidence, appropriate controls, symmetric leakage screening, and validated scorers, stronger contamination claims are not warranted. The results do not establish training-data membership, contamination prevalence, or benchmark-induced score inflation.