Skip to content
AI.info

Research

On-Policy Visual Evidence Distillation

Overview Research area: multimodal large language models and visual agents, specifically on-policy distillation (OPD) for vision-language models that interleave reasoning with image operations. Techni

On-Policy Visual Evidence Distillation
arXiv
2609.36838
Published
2026-09-29
Authors
Shaohang Wei, Feifan Song, Guangyue Peng, Wenhao Yu, Wei Li, Wen Luo, Yang Xu, Yufan Shen, Luke Mao, Yang Du, Asher Qin, Houfeng Wang

AI summary

Overview

Research area: multimodal large language models and visual agents, specifically on-policy distillation (OPD) for vision-language models that interleave reasoning with image operations.

Technical level: Advanced. The paper builds on reverse-KL distillation objectives, token-level gradient analysis, and tool-using multimodal agent training loops.

Scope: The paper introduces ReVuE (Reflection on Visual Evidence), a method that turns a student's own sampled visual interaction trajectories into structured evidence reflections, uses them as training-time context for a stronger teacher, and reweights the distillation loss toward the token positions those reflections most affect.

What This Paper Is About

Visual agents answer questions by repeatedly cropping, zooming, or otherwise manipulating an image and reasoning over what comes back. Because each image operation changes what evidence is available later, a mistake at one stage, such as looking at the wrong region, misreading a correct observation, or mapping a correct fact to the wrong answer, can cascade into a wrong final answer.

Existing multimodal on-policy distillation methods strengthen supervision by building or contrasting auxiliary views of the original image, but they do not explicitly model the link between the student's actions, the observations those actions return, and the student's subsequent reasoning. ReVuE's goal is to use the student's own trajectories to diagnose the first failure in the evidence chain and deliver corrections at the token positions where the teacher's judgment actually changes.

Key Contributions

  1. Visual evidence reflection across student trajectories. ReVuE compares multiple successful student attempts on the same query to identify minimal sufficient evidence and to summarize successful reading and grounding strategies. It also locates the first failure among the stages Acquire, Read, and Ground, producing training-time guidance for a strong teacher.

  2. Reflection-aware grouped reweighting for distillation. Tokens are partitioned according to how strongly the reflection changes the teacher's predictions, and group-wise normalization plus reweighting focuses the distillation updates on high-impact positions.

  3. Cross-model gains and validation of high-impact supervision. Across two model families and 11 benchmarks, ReVuE improves task and tool-call accuracy over OPD baselines while reducing reasoning and tool-call redundancy. On Qwen2.5-VL-7B, masking high-impact positions in the loss reduces performance more than masking equally many low-impact positions.

  4. Critic reliability validation. The stage diagnoses used to build reflections are checked against independent human annotation and against other judge models.

Main Findings

  • Gains across both model families. On 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 families, both ReVuE students outperform all evaluated OPD baselines and RFT in all three weighted category averages (perception, math, general). Perception gains over Vanilla OPD are 2.40% for Qwen2.5-VL-7B and 1.70% for InternVL3.5-4B-Instruct.

  • Benchmark-specific improvements (Qwen2.5-VL-7B). Gains on HRBench 8K (+3.50%) and TreeBench (+3.10%) cover high-resolution understanding and object relations. On InternVL3.5-4B-Instruct, average gains are 2.83% for math and 1.65% for general tasks.

  • Beats answer privilege and visual-privilege baselines. ReVuE exceeds GT-Privileged (which gives the teacher the correct answer) on all 11 benchmarks for both students, and achieves higher category averages than baselines using privileged visual information or visual contrasts.

  • Surpasses its own teacher models. After reflection-assisted distillation, both students also surpass their RL Expert teachers in all three category averages, and exceed GPT-4o in all three averages.

  • Shorter, more accurate, less tool-hungry behavior. On HRBench 8K, ReVuE reduces mean response length from the RL Expert's 458 tokens to 382 while raising accuracy from 71.75% to 74.00% (+2.25%). On TreeBench, it reduces response length by 46.09% and the tool-use rate from 64.7% to 15.8%, while improving accuracy by 3.10% over Vanilla OPD. On V* Bench, VAD uses tools slightly less often (73.3% vs. 74.3%) yet ReVuE achieves 5.24% higher overall accuracy. On VisualProbe, it still uses tools on 83.7% of samples and also achieves higher accuracy.

  • Both components contribute. In the cumulative ablation on Qwen2.5-VL-7B, adding reflection raises all three averages (perception 62.61 to 63.70, math 47.67 to 47.78, general 53.28 to 54.42). Adding impact reweighting raises perception and math by another 1.31% and 0.71% (to 65.01 and 48.49), while the general average drops 0.28% (to 54.14) relative to reflection alone.

  • Impact is highly sparse. Approximately 20% of scored positions in pooled early-training data account for 97.6% of total impact, which motivates concentrating weight rather than averaging uniformly over tokens.

  • Impact ranking beats random selection. Compared with random selection at the same group fractions, ReVuE gains 1.37% in perception and 1.07% in math, but only 0.04% on general tasks. Masking results order consistently as Mask low (65.18 perception) > Mask random (63.85) > Mask top (62.26), with the same ordering on all 11 benchmarks.

  • 20% is the best tested concentration. With λ = 0.5, the 20% setting gives the highest perception (65.01) and math (48.49) averages; 5% and 10% do not exceed it. The general-task average peaks at the 100% setting (54.42), 0.28% above the 20% setting.

  • Diagnoses agree with humans. Final stage labels match 93 of 96 human labels (96.88%, Cohen's κ = 0.957), and agreement remains 95.31% on the 64 incorrect rollouts. Across 768 rollouts, valid pairs agree at 99.2% for a critic rerun and 92.8–93.2% across models.

Methodology in Plain English

ReVuE starts with how a visual agent already behaves. For each training query, the student samples several trajectories (M rollouts per query). Each trajectory is a sequence of actions, executable code, tool-returned text and images, and a final answer. The method then reasons about the whole group in three stages.

First, a critic (in the experiments, Qwen3.5-397B-A17B) compares the trajectories and their observed images. It builds two things for each trajectory. The Anchor records the reference for correct evidence use: the target visual information, the supporting images actually observed within the group, the correct visual fact, and the rule that maps that fact to the answer. The Break Point describes the failure of an incorrect trajectory: which stage failed first in the dependency order Acquire, Read, Ground, and what the specific discrepancy was. This is the visual-evidence reflection.

Second, the method measures how much that reflection matters. The same frozen teacher scores each student trajectory twice: once with only the student's original history, and once with the reflection added as extra training-time context. The student never sees the reflection at inference; only the teacher does at training time. The difference in the log-probability the teacher assigns to the token the student actually generated defines the token impact. A lemma in the paper shows that the change in the distillation gradient is a sum over candidate tokens weighted by these teacher-score shifts, which is why the impact of the sampled token is a meaningful signal. Taking the absolute value captures both reinforcement (teacher support rises) and correction (teacher support falls).

Third, token impact drives the loss. Within each trajectory, the top α fraction of supervised positions by impact form the High group and the rest form the Low group. The total loss weight is split between groups with a parameter λ (set to 0.5, half to each group) and distributed uniformly inside each group. This means each high-impact position receives a larger gradient multiplier than it would under uniform token averaging. If impacts cannot be computed, or if a group is empty, the method falls back to uniform averaging with teacher targets from the unreflected teacher. Everything else is standard on-policy distillation: reverse-KL between the student's distribution and the teacher's distribution, restricted to the teacher's top-K support and renormalized over that support.

The training setup uses Thyme-SFT cold-start checkpoints of Qwen2.5-VL-7B and InternVL3.5-4B-Instruct as students, reproduced Thyme-RL experts as teachers, and Qwen3.5-397B-A17B as the critic.

Why This Matters

Research impact. The paper reframes multimodal distillation as a question of where in the evidence chain supervision should be applied, rather than merely how to present more visual information to the teacher. It also provides a concrete account of how reflection changes the distillation gradient (Lemma 1), and empirical evidence that token-level impact is extremely sparse (20% of positions carry 97.6% of impact). The stage taxonomy Acquire, Read, Ground gives a vocabulary for diagnosing visual-agent failures that other work can build on.

Real-world applications.

  • High-resolution image understanding, where models must locate small task-relevant regions before reasoning.
  • Chart analysis and document question answering, where misreading a value or misapplying a rule produces confidently wrong answers.
  • Agentic assistants that call image tools, where the results show fewer tool calls (TreeBench tool-use rate falling from 64.7% to 15.8%) and shorter responses at higher accuracy.
  • Training smaller, cheaper student models to inherit the visual tool-use behavior of much larger experts, with both ReVuE students surpassing their RL Expert teachers in all three category averages.

Industry relevance. The method targets precisely the pain point of deploying visual agents: tool calls and long reasoning chains are expensive at inference. Reducing response length (458 to 382 tokens on HRBench 8K) and tool-use rate while improving accuracy directly lowers serving cost. The fact that the method works on two different model families (Qwen2.5-VL and InternVL3.5) and beats answer-privileged supervision suggests it is a transferable training recipe rather than a model-specific trick.

Future Directions

  • Broader model coverage. The paper states explicitly that the study focuses on two student model families and image-based question answering, and that broader model coverage remains to be explored.

  • More diverse visual interactions. The current scope is image-based question answering; extending reflection-based supervision to other interaction types (the related work mentions executable code, web search, and longer visual search) is left open.

  • Task-dependent weighting. The general-task average prefers the 100% setting (0.28% above the 20% setting), and Mask low gains 0.17% in perception over the full method while slightly reducing math and general averages. Reconciling these task-dependent preferences, rather than using a single α and λ, is a natural next step.

  • Scaling and cost of the critic. Reflections depend on a separate critic model (Qwen3.5-397B-A17B in these experiments) comparing multiple trajectories per query. Questions remain about how reflection quality and cost scale with critic strength, rollout count M, and whether lighter critics can preserve the 96.88% human agreement and 92.8–93.2% cross-model consistency reported here.

Target Audience

Researchers and engineers working on multimodal large language models, vision-language agents, and knowledge distillation, particularly those training tool-using models where image operations change the available evidence. The paper is also relevant to practitioners who deploy visual agents and care about inference cost, since the reported gains in accuracy come alongside shorter responses and fewer tool calls. Readers should be comfortable with KL-divergence-based distillation and token-level language model training objectives; the method section is mathematically dense, while the experiments and analysis sections are accessible to a broader applied audience.

Authors’ abstract

Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at https://github.com/sylvain-wei/ReVuE

Read the original paper