Skip to content
AI.info

Research

Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction

Overview Research area: Efficient multimodal large language models (MLLMs) — specifically, training-based visual token reduction for extreme compression regimes. Technical level: Intermediate. The pap

Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction
arXiv
2609.32353
Published
2026-09-26
Authors
Junxian Li, Ruixuan Yang, Tianao Zhang, Tiange Xu, Weisheng Dong, Yulun Zhang

AI summary

Overview

Research area: Efficient multimodal large language models (MLLMs) — specifically, training-based visual token reduction for extreme compression regimes.

Technical level: Intermediate. The paper assumes familiarity with transformer-based vision-language models, token pruning, knowledge distillation, and KL/Jensen-Shannon divergence, though the core idea is explained conceptually.

Scope: This paper proposes LT-OPD, a training framework that combines on-policy self-distillation with a budget-level curriculum to help a heavily compressed MLLM recover capability lost when only a small fraction of visual tokens is retained.

What This Paper Is About

Multimodal large language models process images as large numbers of visual tokens, which makes inference expensive. Existing methods try to cut these tokens down or fine-tune models to cope with fewer of them, but performance collapses when the token budget is extremely low (for example, keeping only 5%). The authors argue that compression changes not just how much visual evidence a model has, but also the sequence of generation states it visits at inference time — and that training on fixed response prefixes ignores this. Their goal is to train the compressed model on the trajectories it actually produces, using its own full-token version as a teacher.

Key Contributions

  1. A reframing of extreme visual-token reduction as a training problem. The authors highlight that severe compression removes visual evidence and also shifts the generation states the compressed model encounters, motivating adaptation on the model's own generated trajectories rather than fixed prefixes.

  2. LT-OPD (Low-Token On-Policy Distillation). A framework where a low-token student rolls out its own responses and a frozen full-token copy of the same MLLM supplies distributional supervision at the states the student visits, using Jensen-Shannon divergence.

  3. A budget-level curriculum. A three-stage schedule that starts training at a larger token budget and progressively reduces it to the target budget (5%), stabilizing on-policy learning when visual evidence is severely limited.

  4. Extensive evaluation. Tests across nine benchmarks, two model scales (4B and 9B), and three MLLM families (Qwen3.5, GLM-4.6V, LLaVA-OV), plus efficiency measurements and comparisons against reinforcement-learning and training-based baselines.

Main Findings

  • Large retained-performance gain under 5% tokens: On Qwen3.5-4B, LT-OPD raises average retained performance from 68.6% to 82.3% while preserving only 5% of visual tokens.

  • Beats training-free and training-based baselines at the same budget: The strongest training-free baseline, DivPrune, reaches 77.6% (4.7 percentage points below LT-OPD), and the training-based baseline EPIC reaches 73.7% (8.6 points below).

  • Broad rather than benchmark-specific improvement: LT-OPD achieves the best result among prior methods keeping 5% of visual tokens on 8 out of 9 benchmarks, including all categories. Clear gains appear on V*Bench (74.9 vs. 70.7), MME (2,077.0 vs. 1,915.4), and OCRBench (535 vs. 477).

  • Better than training-free methods with more tokens: With only 5% of visual tokens, LT-OPD still exceeds all training-free methods retaining 10% tokens in average retain ratio (82.3% vs. at most 80.2%), and nearly matches the best results at 15% or 20% budgets.

  • Both components matter (ablations): Curriculum learning raises the average retain ratio of plain SFT by 2.7 percentage points (70.4% to 73.1%); OPD-only training reaches 81.5%; curriculum with fixed-prefix JSD distillation reaches 77.6%; the full LT-OPD reaches 82.3%.

  • Gains transfer across model scales: On Qwen3.5-9B, average retained performance improves from 71.1% to 81.4%.

  • Gains transfer across MLLM families: GLM-4.6V-9B improves from 73.8% to 86.3%, and LLaVA-OV-1.5-4B improves from 71.8% to 84.1%.

  • Efficiency is preserved: With 5% visual tokens, LT-OPD reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% with no additional inference overhead. It matches HiPrune's 0.146 FLOPs ratio while improving retained performance by 21.9 percentage points, and requires 27.0% fewer FLOPs than EPIC (0.146 vs. 0.200).

  • Modest latency improvement: End-to-end latency improves from 1331.5 ms to 1293.7 ms (1.029x), consistent with the observation that aggressive visual-token reduction mainly benefits prefill, while autoregressive decoding dominates per-token computation and memory access.

  • Better than general-purpose RL: GRPO reaches 75.8%, GSPO 76.8%, and DAPO 78.2% average retain ratio, all below LT-OPD's 82.3%.

  • Robust to teacher regularization choice: Using the frozen initial policy (82.3%), trust-region regularization (81.5%), or EMA (81.3%) produces scores that do not change much.

  • Superior compression-gap recovery: On the recovery metric 100*(S_method − S_pruned)/(S_full − S_pruned), LT-OPD recovers the most across four evaluated benchmarks and exceeds RL methods by over 45 points on V*Bench.

  • Multi-image results: On MuirBench and MVBench, LT-OPD outperforms its training-free counterpart (CDPruner) and EPIC by over 12% in average retain ratio.

  • Training cost: Training used 14,000 questions, 175 optimizer updates, and 112,000 generated trajectories, running 15.1 hours on eight GPUs — within the 12.2–48 hours reported for EPIC, LLaVA-Mini, and similar methods on the same devices.

Methodology in Plain English

The setup uses a single pretrained MLLM twice. One copy is frozen and keeps all visual tokens, acting as the teacher. The other copy is the student and receives only a small fraction of visual tokens (target 5%).

Instead of training the student on fixed example responses, the student first generates its own responses from its compressed visual input. Those generated responses are treated as fixed (no gradients flow through the sampling process). Then, at each position in the student's own generated text, both the teacher and the student predict the next token — the teacher conditioned on the full image and the student on the compressed image. The training objective pulls these two next-token distributions together using Jensen-Shannon divergence, with a 0.5 mixture coefficient. Because the supervision happens on prefixes the student itself produced, the student learns to correct itself exactly in the situations it will face at inference. The teacher never generates its own separate trajectory, and only the compressed student is kept for inference.

To keep training stable, the authors introduce a three-stage curriculum over the visual-token keep ratio. Stage one (steps 0 to 13) uses a higher budget of (2α+1)β with α set to 2. Stage two (steps 14 to 99) follows a cosine schedule that smoothly reduces the budget. Stage three (step 100 onward until the total training steps) uses the final target ratio β. This lets the student first build useful trajectories with more visual evidence before adapting to the hardest condition.

Implementation details: the token selection uses CDPruner adapted to Qwen3.5, applied after the native visual merger and before the language model, parameter-free and run in FP32 without gradients. Distillation support at each position is the union of the student's top 100 tokens, the teacher's top 100 tokens, and the sampled token, with remaining probability mass aggregated into one tail category. Training data, called LT-14K, contains 14,000 single-image questions drawn from LLaVA-Instruct-665K, OneThinker, PixMo, and Vision-OPD sources. Full-parameter updates cover the language model, vision encoder, and native visual merger. Training used 8 NVIDIA A100 80GB PCIe GPUs, AdamW, a learning rate of 1×10⁻⁶ for the language model and merger and 2×10⁻⁷ for the vision encoder, a 14-update warmup followed by a constant schedule, FSDP across 8 data-parallel ranks, and a maximum response length of 1,024 new tokens.

Why This Matters

Impact on research: The paper argues that under extreme compression, token selection and model adaptation should be considered jointly rather than in isolation. It also shows that on-policy self-distillation — previously explored mainly for language models or for enhancing policy performance with stronger teachers — can serve a different purpose: recovering capability lost to input compression. The result that a 5%-token model can match or exceed training-free methods using 10% tokens suggests that training signal, not token ranking, becomes the dominant bottleneck at very low budgets.

Real-world applications:

  • On-device visual assistants and phone-based multimodal chatbots, where memory and prefill compute are tightly constrained.
  • Edge robotics and embedded vision systems that need to interpret camera input with limited hardware.
  • High-throughput document and OCR processing pipelines, where the paper reports gains on OCRBench and TextVQA.
  • Long-context multimodal systems handling many images, where KV-cache and prefill costs compound (the paper reports 85.2% KV-cache reduction).

Industry relevance: The 85.2% KV-cache and 85.4% prefill FLOP reductions come with no additional inference overhead, meaning deployments can adopt the compressed student directly. The framework is model-agnostic in the authors' tests, working on Qwen3.5, GLM-4.6V, and LLaVA-OV backbones, and its 15.1-hour, 14K-question training budget is far below typical large-scale instruction-tuning runs.

Future Directions

  • Video and multi-image understanding: The authors note strong results on MuirBench and MVBench and suggest LT-OPD may generalize to video understanding tasks, which they identify as a future direction.

  • Joint optimization of selection and adaptation: The conclusion explicitly calls for treating effective token selection and effective adaptation as a combined problem rather than separate ones.

  • Training efficiency and cost: Although 15.1 hours on eight GPUs is described as acceptable, further reducing the cost of on-policy rollouts (8 rollouts per question over 175 updates) is a natural open problem.

  • Objective design trade-offs: The authors report that while JSD is better on perception-heavy benchmarks, forward/reverse KL achieve better scores on OCR benchmarks, leaving open which objectives are best for which task families.

Target Audience

Researchers and engineers working on MLLM inference efficiency, visual token pruning and merging, and knowledge distillation for multimodal models. It is also relevant to practitioners deploying vision-language models under tight memory or compute budgets, and to readers interested in how on-policy training signals transfer to input-compression problems. Basic familiarity with transformer architectures and distillation losses is assumed.

Authors’ abstract

Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.

Read the original paper