Research
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models
Overview Research area: Efficient inference for multimodal large language models (MLLMs), sitting at the intersection of computer vision, vision-language modeling, and model compression. Technical lev

- arXiv
- 2609.34972
- Published
- 2026-09-28
- Authors
- Jingdi lei, Junxian Li, Di Zhang, Zhanqiu Zhang, Yiwen Guo, Soujanya Poria
AI summary
Overview
Research area: Efficient inference for multimodal large language models (MLLMs), sitting at the intersection of computer vision, vision-language modeling, and model compression.
Technical level: Advanced. The paper assumes familiarity with Transformer internals (queries, keys, values, hidden states, feed-forward layers), low-rank decomposition, attention masking, and knowledge distillation.
Scope: The paper proposes a method called δ-Vision that replaces the repeated Transformer evolution of visual tokens with lightweight low-rank MLP adapters that predict layer-wise visual memory, evaluated on image, multi-image, and video benchmarks across several MLLM backbones.
What This Paper Is About
Multimodal LLMs pay a large computational cost because long visual token sequences are re-processed at every Transformer layer of the language model. Prior work attacks this by pruning visual tokens, but pruning is irreversible: any visual evidence discarded early cannot be recovered by later layers. This paper asks the opposite question: can all visual tokens be kept while making their layer-by-layer evolution much cheaper?
Key Contributions
-
Empirical evidence of hidden-channel redundancy and predictability. The authors show that visual information influences the text stream through a low-dimensional subspace, and that layer-specific visual states are highly predictable from the initial visual embeddings by lightweight MLPs.
-
δ-Vision, a visual-memory prediction framework. Repeated Transformer updates of visual tokens are replaced by low-rank adapters that construct layer-wise visual memory, while the language-model backbone stays frozen and every visual token remains available as key-value context for text retrieval.
-
Two adapter variants. An embedding adapter that predicts each layer's memory directly from the initial visual embeddings, and a recurrent adapter that progressively updates the memory across layers from the preceding one.
-
Broad empirical validation. Experiments across multiple MLLM backbones and single-image, multi-image, and video benchmarks show stronger accuracy-efficiency trade-offs than visual token pruning under comparable computation.
Main Findings
-
Low-rank sufficiency of visual effects. When visual-to-text attention is blocked and only a few restored directions are reintroduced, most of the lost accuracy returns. On Qwen3-VL-4B, blocking visual information in all layers reduces RWQA from 71.2 to 46.1 and MMStar from 64.9 to 25.6, while restoring only 128 hidden directions recovers them to 70.7 and 64.6.
-
Concentrated spectral structure. Visual queries, per-head QK^T matrices, and visual attention outputs have strongly concentrated spectra. Visual attention outputs need only 38.0-103.5 directions on Qwen3-VL-4B and 29.6-37.6 directions on LLaVA-1.5-7B to retain 95% of squared singular-value energy.
-
Accuracy recovery from low rank. On Qwen3-VL-4B, rank 128 reaches 92.2 on SQA, 69.9 on RWQA, and 61.0 on MMStar, compared with 93.4, 71.4, and 64.8 when the full visual effect is retained.
-
Main single-image result. On Qwen3-VL-4B, δ-Vision reaches an average score of 74.4 with the embedding adapter, retaining 92.9% of the uncompressed model performance (80.1). This is 9.1 points above the strongest 5%-retention pruning baseline, DivPrune (65.3), and 9.4 and 7.5 points above the training-based baselines LLaVA-Mini (65.0) and EPIC (66.9). It also slightly surpasses the best result among pruning methods retaining 20% of visual tokens.
-
Recurrent adapter adds further gains. The recurrent variant improves the average from 74.4 to 75.4, indicating that lightweight cross-layer evolution helps beyond the embedding-only prediction.
-
Efficiency on video. On Video-MME, δ-Vision uses 17.90% of uncompressed model FLOPs, comparable to the 16.58-23.90% range of methods retaining 5% of visual tokens, and achieves 1.30x end-to-end and 1.50x prefill speedup. Reported peak memory is 12.97 GB versus 12.94 GB for the uncompressed model, so memory is not reduced.
-
Depth matters non-uniformly. On Qwen3-VL-4B, blocking visual information in the middle layers (13-22) causes the largest degradation, while first and last layers are largely insensitive. LLaVA-1.5-7B shows a more distributed, task-dependent pattern, with early layers more important on MMStar.
-
Cross-backbone generalization. δ-Vision scores 78.6 on Qwen3-VL-30B-A3B (versus 71.3 for DivPrune at 5% retention, 82.7 vanilla), 71.6 on Qwen3.5-4B (versus 62.7, 74.7 vanilla), and 61.9 on LLaVA-1.5-7B (versus 59.8, 65.6 vanilla), retaining roughly 94-96% of vanilla performance across the three backbones.
-
Multi-image and video training helps. The embedding adapter trained only on single-image data averages 49.5 on MuirBench, Video-MME, and MVBench; adding multi-image and video training data improves this to 52.0, with the largest gains on MuirBench (40.7 to 45.5) and MVBench (57.4 to 59.5). This surpasses the strongest 5%-retention pruning baseline at 50.2.
-
Distillation objective. Adapters are trained with Supervised-KD using forward KL divergence between teacher and student predictive distributions over answer tokens; the appendix reports this beats standard SFT and on-policy distillation while avoiding rollout cost.
Methodology in Plain English
The authors start with a diagnostic experiment. They take a trained MLLM, keep all visual tokens, and project the visual hidden states onto a much smaller subspace before letting subsequent computation continue. Accuracy mostly survives, which suggests the visual information actually used is low-dimensional. They then train a small MLP per layer to predict each layer's visual state from the initial visual embeddings, and find high cosine similarity and low reconstruction error.
Building on this, δ-Vision keeps the vision encoder, the multimodal projector, and the language model frozen, and trains only small low-rank adapters (bottleneck rank 128 by default). At each layer, an adapter constructs a "visual memory" as a residual correction to the initial embeddings (embedding adapter) or to the previous layer's memory (recurrent adapter). That memory is pushed through the original frozen layer norm and key/value projections, so text queries can still attend to every visual token. What is removed is the visual side of the computation: no visual queries, no visual attention outputs, no visual feed-forward updates, and no self-attention among visual tokens. Visual memories are used once as read-only context and then discarded, rather than being propagated as Transformer hidden states.
Training uses the original MLLM as a frozen teacher. The single-image configuration trains on allenai/pixmo-ask-model-anything for 2,000 steps on 8 GPUs with global batch size 32; the multi-image and video configuration uses 20,000 samples from allenai/Molmo2-MultiImageQA, 44,000 from lmms-lab/M4-Instruct-Data, and 64,000 from Video-R1/Video-R1-data for 4,000 steps. AdamW with learning rate 5×10⁻⁵, weight decay 0.01, gradient clipping 1.0, cosine schedule with 3% warmup, and seed 44 are used throughout. DeepStack is disabled in Qwen3-VL for all compared methods to keep the comparison fair.
Why This Matters
The paper argues that multimodal inference efficiency has more than one axis. Token pruning shortens the sequence; δ-Vision instead changes how visual states are built, keeping the full visual evidence accessible. The paper states that the two mechanisms are complementary and reports in the appendix that δ-Vision can be combined with DART and DivPrune.
- High-resolution image question answering and captioning, where long visual token sequences dominate cost.
- Video understanding, the setting where the reported 17.90% FLOPs, 1.30x total speedup, and 1.50x prefill speedup were measured on Video-MME.
- Multi-image reasoning tasks such as interleaved document, screenshot, or diagram understanding, evaluated on MuirBench.
- Latency-sensitive or resource-constrained deployment of large multimodal models, where prefill cost is the bottleneck.
Industry relevance: The method trains only small adapters against a frozen backbone and reuses the original key/value projections, which fits existing serving stacks without retraining the base model. The reported speedup is a prefill-oriented gain, which matters for long-context multimodal serving, though the paper does not report a memory reduction (peak memory was 12.97 GB versus 12.94 GB for the uncompressed baseline).
Future Directions
- Layer-selective skipping. The non-uniform depth sensitivity found in the intervention study suggests that per-layer decisions about how much visual computation to run could be exploited further; the paper says this is validated in its appendix on flexible layer-wise skipping.
- Combining with token pruning. The authors state the two efficiency axes are complementary and report preliminary combined results with DART and DivPrune, leaving fuller joint designs open.
- Architectures beyond standard attention. Qwen3.5-4B uses a hybrid attention architecture, which the paper says receives additional analysis in its appendix but is not developed into a general design rule.
- Broader backbone coverage. The paper evaluates additional models including LLaVA-1.5-13B, LLaVA-v1.6-Mistral-7B, and Qwen3-VL-8B in its appendix, and the generality of visual-memory prediction across other modality and architecture families remains an open question.
Target Audience
Researchers and engineers working on efficient multimodal LLM inference, visual token compression, and knowledge distillation for vision-language models. It is also relevant to practitioners deploying video or high-resolution image models under latency constraints, and to readers interested in the dimensional structure of visual representations inside Transformers. The paper's analysis sections assume comfort with spectral statistics (singular values, effective rank) and attention mechanics.
Authors’ abstract
Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose $δ$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, $δ$-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.