Skip to content
AI.info

Research

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

Overview Research area: Artificial intelligence — specifically agentic AI, test-time verification, and long-horizon agent evaluation. Technical level: Advanced. The paper assumes familiarity with LLM

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
arXiv
2610.00972
Published
2026-10-01
Authors
Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee

AI summary

Overview

Research area: Artificial intelligence — specifically agentic AI, test-time verification, and long-horizon agent evaluation.

Technical level: Advanced. The paper assumes familiarity with LLM agents, rollout sampling, test-time scaling, LLM-as-a-judge methods, and information-theoretic measures such as entropy.

Scope: The paper introduces VeriHarness, a training-free, plug-and-play harness that lets the same frozen model that generated an artifact also verify it by testing claims against the task environment, and evaluates it across five long-horizon workspace benchmarks and two frontier models.

What This Paper Is About

Long-horizon AI agents increasingly produce reports, spreadsheets, code patches, and other multi-file artifacts whose correctness is hard to assess, and even frontier models struggle to produce reliable ones. The paper asks how verification can be strengthened when the verifier is the same model as the generator, with no access to reference answers or grading rubrics at test time — a setting the authors note reflects practice at the frontier, where no stronger judge exists. The goal is a general-purpose verification harness that inspects many sampled rollouts, checks their competing and shared claims against environmental evidence, and returns a better artifact rather than just a score.

Key Contributions

  1. A verification principle. The paper identifies two complementary checking tasks in same-model verification: resolving disputed claims (where rollouts disagree and correct alternatives may be exposed) and challenging consensus claims (where rollouts agree but may share an error), both settled against evidence from the environment.

  2. A general-purpose harness. VeriHarness is described as the first agentic verification harness for long-horizon tasks. It is training-free and plug-and-play across benchmarks and models, and is evaluated both for selecting a rollout from the pool and for revising it.

  3. Evolving verification skills. The authors show that a fixed model can accumulate useful verification skills from failure feedback on development tasks, starting from either an empty library or a human-authored one, without changing the model, protocol, or general tools.

  4. A released rollout pool. The authors release approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.

Main Findings

  • VeriHarness leads all evaluated selection baselines. On the primary selection comparison it exceeds every baseline in all ten model–benchmark settings and improves the single-rollout average by 4.4 points with Gemini 3.5 Flash and 4.1 points with Claude Opus 4.8 (Table 2).

  • Revision on top of selection adds substantial gains. Applying the evidence-backed revision plan improves all ten model–benchmark settings and raises the average gain over a single rollout to 6.2 points with Flash and 6.4 points with Opus. Per benchmark with Flash the gains are +6.7 (APEX-Agents), +8.9 (Workspace-Bench Lite), +5.0 (WorkBuddy Bench), +4.9 (SpreadsheetBench 2), and +5.3 (JobBench); with Opus they are +11.7, +6.8, +4.7, +4.2, and +4.8 respectively.

  • Consensus hides errors; disagreement reveals correct alternatives. On ten-rollout pools of Claude Opus 4.8 on APEX-Agents, 34% of consensus values are judged incorrect, while 74% of disputed claims contain a correct candidate and the most frequent value is correct in only 47% of cases (Figure 2).

  • Environment access alone is not enough. An agentic verifier given the same workspace and evidence tools but no VeriHarness protocol or skills recovers only about half of the harness's gain. Baseline methods that only read the pool — majority voting, best-of-N judging, pairwise tournament, LLM-as-a-Verifier — yield only modest average gains on these long-horizon tasks.

  • The resolver and challenger recover different errors. With Claude Opus 4.8, the disagreement resolver alone raises the five-benchmark average from 49.6 to 54.3 and the consensus challenger alone to 52.6, while the two together reach 56.1, so their gains are complementary. Merging the two investigations into one context costs 0.8 points.

  • Selection gains concentrate in pools that disagree. On the 154 APEX-Agents consensus pools, the harness's selection scores only 0.6 points above the pool mean; on the 262 disputed pools it scores 6.1 points above. Within the disputed pools the gain grows with entropy, from 4.5 points on low-entropy pools to 7.7 on medium-entropy pools (Figure 5).

  • Challenging consensus is generally safe. On APEX-Agents the challenger never refuted a claim that every rollout had right. In about 70% of the shared errors it left standing, it had checked an intermediate result that was correct while the error sat in a later step — a blind spot the authors attribute to the verifier tending to stop where the rollouts stopped.

  • Skills evolved from an empty library beat the human-authored library on held-out tasks. Library C (evolved from empty) improves the held-out score by 11.0 points on APEX-Agents and 5.7 points on SpreadsheetBench 2 relative to an empty library A, and also exceeds the human-authored library B. Evolving from B (library D) yields the highest final score on both benchmarks, adding 6.8 points on APEX-Agents and 3.7 points on SpreadsheetBench 2 over B.

  • Evolved skills are more specific. Of the 95 checks in the four evolved libraries, 5 restate a human check, 48 make a human check concrete with a script, a constant, or a document type, and 42 are new.

  • The protocol transfers to existing agent harnesses. Running it inside Gemini CLI, Claude Code, or Codex preserves most of the selection improvement, though with Opus those CLIs fall back to the single-rollout level on JobBench, as does Claude Code on SpreadsheetBench 2.

Methodology in Plain English

The researchers start from a task and a frozen pool of ten rollouts per task per model, all produced by the same base model (Gemini 3.5 Flash or Claude Opus 4.8 at thinking level high) that also acts as the verifier. Every method compared receives the same pool at equal generation cost.

They examine artifacts at the level of claims — individual assertions about values, interpretations, or requirement satisfaction. Claims shared unanimously across all rollouts are called consensus claims (entropy zero); claims where rollouts differ are disputed claims (entropy above zero).

VeriHarness then runs three stages in separate model contexts:

  1. A disagreement resolver compares artifacts, traces competing values to their sources, and runs checks that best distinguish them — for example verifying which version of a document supersedes another.
  2. A consensus challenger proposes ways a universally shared value could be wrong and tests them against source metadata, file conventions, or the task description, including requirements that every rollout overlooked.
  3. Adjudication in a fresh context reviews both evidence records and produces a base rollout, a revision plan, and a list of unresolved claims, which are then applied to deliver the final artifact plus a verification record.

The harness is built from a workspace (artifacts, final states, and traces exposed as files), evidence tools (reading, recomputation, code execution), a verification protocol, and a skill library of reusable failure modes and checking procedures. Finally, the authors test self-evolution: they split each benchmark about 3:1 into development and held-out sets, let the verifier read development failures with grader feedback and propose candidate libraries, keep the candidate with the highest development score, and record held-out scores without ever using them for selection.

Why This Matters

Impact on research. The paper argues that verification capability "does not have to wait for a stronger model" — it can be built from the structure of a model's own rollouts, the evidence in its environment, and accumulated experience. It also connects verification to training: the verification record links every claim to its check, evidence, and verdict, giving a claim-level signal that could supervise the generator. The release of roughly 26,000 rollouts at a cost of over $100,000 is offered as infrastructure for future work.

Real-world applications:

  • Verifying professional deliverables such as consulting memos, financial models, and legal analyses, where claims must be traced to sources and quantities recomputed (the APEX-Agents setting).
  • Auditing spreadsheets, including catching shifted ranges in growth-rate formulas and unit or sign-convention errors (SpreadsheetBench 2).
  • Checking multi-file digital work products across white-collar occupations (JobBench) and file-heavy workspaces (Workspace-Bench Lite).
  • Reviewing code patches and office/web artifacts before delivery (WorkBuddy Bench).

Industry relevance. The system deploys inside existing agent CLIs (Gemini CLI, Claude Code, Codex) and gains most of its improvement there, suggesting the protocol is not tied to one vendor implementation. It also fits deployment settings where quality matters more than latency, such as professional deliverables, and it uses the same model the generator already uses rather than requiring a stronger, more expensive judge.

Future Directions

  • From records to reward. The paper does not assign a scalar score to individual rollouts; turning claim-level verification records into a calibrated rollout-level score or reward signal is explicitly left to future work.
  • Learning where to look. The challenger's blind spot — stopping where the rollouts stopped and checking correct intermediate results while missing later errors — points to a need for more prior knowledge of how artifacts fail.
  • Beyond the same-model setting. The authors restrict the study to a verifier identical to the generator; measuring whether a stronger verifier helps, and separating that benefit from the capability gap, is out of scope here.
  • Cost and latency trade-offs. Verification uses a pool of ten rollouts per task and adds wall-clock time through multi-turn investigation, so reducing pool and verification cost remains open.

Target Audience

Researchers and engineers working on LLM agents, test-time scaling, and automated evaluation; practitioners building verification or quality-control pipelines for long-horizon deliverables such as reports, spreadsheets, and patches; and anyone studying self-improving agent harnesses or skill-library evolution without weight updates. Readers without background in agent rollouts, majority voting, and LLM-as-a-judge comparisons will find the methodology sections demanding.

Authors’ abstract

As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.

Read the original paper