Skip to content
AI.info

Research

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

Overview Research area: Embodied AI / agentic self-improvement — specifically automated verification (judging) of vision-language-action policies for autonomous driving and robot navigation. Technical

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
arXiv
2610.08761
Published
2026-10-06
Authors
Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao, Chensheng Peng, Boyi Li, Thomas Tian, Zheng Lian, Yan Wang, Jiaqi Ma, Boris Ivanovic, Marco Pavone, Wenhao Ding

AI summary

Overview

  • Research area: Embodied AI / agentic self-improvement — specifically automated verification (judging) of vision-language-action policies for autonomous driving and robot navigation.
  • Technical level: Advanced. The paper assumes familiarity with reinforcement fine-tuning (GRPO), supervised fine-tuning, vision-language models, reward hacking, and rubric-based evaluation.
  • Scope in one sentence: The paper introduces VeriFine, an agent-harness framework that lets a policy, its training curriculum, and its verification judge co-evolve across repeated rounds, using selective human guidance to repair the judge when it becomes the bottleneck.

What This Paper Is About

Self-improving AI policies keep producing new kinds of mistakes, but most systems rely on a single fixed judge whose ability to spot errors goes stale as the policy gets better — leading to inaccurate feedback, reward hacking, and diminishing returns. This is especially hard in embodied reasoning, where a trustworthy judge must account for spatial grounding, causal reasoning, and safety-aware decision-making. VeriFine addresses this by treating verification itself as something that must be scaled and improved alongside the policy, rather than held constant.

Key Contributions

  1. VeriFine framework: An agent-harness framework for continuous self-improvement through scaling verification, built around the co-evolution of the policy, the training curriculum, and the judge.
  2. Policy Improvement Loop: A loop that uses judge-identified recurring failure patterns to select high-value data and optimize the policy autonomously, without human intervention.
  3. Judge Improvement Loop: A loop that detects when verification has become the bottleneck, selectively queries human guidance on informative failure cases, and performs "coactive calibration" — a process where humans and agents resolve disagreements rather than treating human labels as infallible ground truth — so the judge tracks the evolving policy frontier.
  4. Demonstrated cross-task, cross-method self-improvement: Experiments on driving (reinforcement fine-tuning) and robot navigation (supervised fine-tuning) show gains in both policy reasoning and judge capability.

Main Findings

  • Driving policy improvement across three rounds: On a fixed reference-based evaluator (VeriFine-Judge-RB) over a 2,772-sample test set, the policy reasoning score rose from 60.56 for the base policy to 71.61 after round R3, an 18.2% relative improvement. Under the final reference-free judge the relative gain was 22.2%.
  • Judge-guided data selection beats random: With VeriFine-Judge-RB held fixed as the reward, judge-guided curriculum selection improved the reasoning score from 64.03 (random selection) to 67.70. Across rounds, judge-guided selection improved reference-based performance over random selection by 4.5–14.6% (reported in the appendix table).
  • Reference-free judge matches reference-based judges: VeriFine-Judge-R1 reached 67.30, comparable to the strongest reference-based configuration at 67.70, without relying on costly annotations during inference. R2 and R3 reached 70.59 and 71.61, with R3 exceeding the strongest reference-based baseline by 5.8%.
  • Judge alignment improves with calibration: Teacher-judge correlation with human scores increased from 0.55 to 0.85 and MAE decreased from 0.25 to 0.13. The final distilled student judge achieved r = 0.82 and MAE = 0.17, comparable to VeriFine-Judge-RB (0.82 / 0.16) and better than LingoJudge (0.65 / 0.23). On the judge test set, VeriFine-Judge-RB achieved r = 0.82 versus 0.65 for LingoJudge.
  • Coactive calibration adds value beyond the backbone model: With the Claude Opus 5 backbone held fixed, coactive calibration improved correlation from 0.67 to 0.85 (appendix figure).
  • The judge expands its verification boundary: The R1 student achieved r = 0.72 on the first-round test subset but only r = 0.07 on the later, more challenging subset; refinement raised the latter to 0.74 in R2. R3 then targeted remaining errors across accumulated scenarios without adding a new test subset.
  • Test-time scaling with the evolved judge: On 401 independent scenarios with six reasoning-action candidates each, the final judge selected outputs with a human-rated reasoning score of 78.6, outperforming the R2 judge and LingoJudge and performing comparably to VeriFine-Judge-RB. Since candidate generation was held fixed, gains reflect better selection.
  • Navigation transfer: On robot navigation using VLNVerse scenarios with a Qwen3-VL 2B policy trained by supervised fine-tuning, the final policy reasoning score improved by 14% from 68.09 to 77.59, and final judge-human correlation approached 0.85.
  • Lowest ADE without action rewards: The final driving policy achieved the lowest ADE even though no action reward was used in training.

Methodology in Plain English

The system runs two coupled loops. In the Policy Improvement Loop, the policy generates outputs over a human-quality-checked anchor evaluation set; a reference-free judge — one that scores reasoning directly from the physical context instead of comparing against a human-written reference — produces an overall score plus structured diagnostic scores. The agent aggregates these to spot recurring failure patterns, turning them into selection criteria over a large candidate data pool. Selection does two jobs: it prioritizes contexts where the policy is weak, and it filters out unreliable targets (a scenario that exposes a weakness is useful for learning, but an incorrect response to it is a bad demonstration). The selected curriculum then drives either reinforcement fine-tuning with GRPO (driving) or supervised fine-tuning (navigation).

When progress plateaus, the framework runs the Judge Improvement Loop. Progress stagnation and a growing gap between judge reward and actual anchor-set performance are monitored as trigger signals. The agent then assembles a set of uncertain, inconsistent, or high-impact cases — drawn from the newly updated policy's outputs so they reflect current failures, and deliberately including both good and bad outputs. Human experts first coarsely calibrate scores; the agent then autonomously proposes, evaluates, and selects rubric revisions. When alignment plateaus, disputed cases are shown to humans with the agent's diagnostics, and humans may confirm, correct dimensions, clarify criteria, or even revise their own prior annotations. This cycle repeats until no alignment gain remains, converging on a shared rubric of physical reasoning. Cumulative evaluation sets are retained so new refinements don't degrade previously acquired capability.

To keep large-scale optimization affordable, the judge uses a teacher-student design: a frontier vision-language model (Claude Opus 5) serves as the rubric teacher, and its capability is distilled into a compact student built on a Qwen3-VL 2B backbone, pretrained on internal driving data with a VLA planning task. Across rounds, each new policy is trained from the same base initialization under a matched per-policy training budget, so comparisons reflect the accumulated supervision state rather than extra compute.

Why This Matters

  • Impact on research: The paper reframes verification from a static component into a scalable, evolving subsystem that must be improved in step with the policy. It also introduces a middle path between fully autonomous self-improvement and expensive human annotation — selective, targeted human guidance at the moment the judge hits its limit.
  • Real-world applications:
    • Autonomous driving, where reasoning about scenes must be judged for safety and instruction consistency without a human-written reference for every clip.
    • Robot navigation, where policies must be scored on spatial grounding, goal understanding, and stopping decisions.
    • Large-scale training-data curation, filtering out well-solved scenarios that no longer provide learning signal.
    • Test-time selection/ranking of multiple candidate plans for embodied agents.
  • Industry relevance: The work comes from NVIDIA with UCLA, UC Berkeley, and Stanford, and is evaluated on an internal dataset of 2 million driving clips across 25 countries using an 8B policy model. The teacher-student distillation makes the verification pipeline cheap enough to run inside large-scale optimization loops, which is the practical prerequisite for deploying self-improving embodied systems.

Future Directions

  • Extending beyond driving and navigation: Whether the two-loop framework transfers to other embodied domains requiring spatial grounding, temporal understanding, causal reasoning, and safety-aware decision-making is not established in the content provided.
  • Reducing reliance on frontier teacher models: The teacher-student design depends on a frontier VLM (Claude Opus 5) for rubric quality; how far this can be pushed with weaker or fully open teachers is not reported.
  • Understanding when human guidance is truly necessary: The paper argues that "little human guidance" suffices when the judge nears its boundary, but the exact thresholds, trigger rules, and check frequency are described as being in the appendix section A.4.3, which is not included in the provided content.
  • Additional analyses promised but not included: The provided content is truncated at Appendix A.1; the referenced appendix tables and figures (judge-guided selection gains of 4.5–14.6%, student architecture ablations, latency performance, and the coactive calibration figure) are not available in the excerpt, so their specific values are not reported here.

Target Audience

Researchers and engineers working on self-improving AI systems, reinforcement learning from AI feedback, reward modeling, and agentic pipelines. It is also relevant to practitioners in autonomous driving and robotics who need scalable evaluation of free-form reasoning, and to those studying human-in-the-loop alignment who want to see how minimal, targeted human input can be used rather than exhaustive annotation.

Authors’ abstract

Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.

Read the original paper