Research
Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
Overview Research area: Multimodal large language models (MLLMs), specifically latent visual reasoning (LVR) — reasoning performed in continuous latent token space rather than with textual chain-of-th

- arXiv
- 2609.34563
- Published
- 2026-09-28
- Authors
- Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li, Linghan Xu, Jiaye Wu, Zheng Zhang, Xiang Xu, Xuanbai Chen, Farhan Tejani, Jakub Zablocki, Julia Xu, Yifan Xing
AI summary
Overview
Research area: Multimodal large language models (MLLMs), specifically latent visual reasoning (LVR) — reasoning performed in continuous latent token space rather than with textual chain-of-thought — and the reinforcement-learning post-training (GRPO) used to train it.
Technical level: Advanced. The paper assumes familiarity with vision-language model architectures, GRPO-style RL post-training, attention readout analysis, and teacher forcing.
Scope (one sentence): The paper diagnoses why latent visual reasoning tokens fail to track answer-relevant image evidence, proposes a training-time fix (ReaLVR) that supervises the model's own free-running latent trajectory, and validates it across five benchmarks and six backbones from three model families up to 235B parameters.
What This Paper Is About
Latent visual reasoning lets a multimodal model do its intermediate computation in continuous hidden states instead of writing out every step in words. The problem is that nobody can see those latent tokens, so it is unclear what they actually learn: the paper's analysis shows they barely react to image edits that change the correct answer. The goal is to give latent tokens direct supervision about where in the trajectory to accept stronger supervision and what visual evidence to preserve, without changing the model architecture or the inference procedure.
Key Contributions
-
Identifies the "latent evidence-credit gap." A behavioral and representation analysis shows that vanilla LVR does not reliably preserve answer-relevant visual evidence in its generated trajectory — latent tokens respond only weakly to image perturbations that alter the correct answer. The authors attribute this to the absence of explicit supervision during GRPO training, where only generated text positions are scored.
-
Introduces ReaLVR. A method that regenerates a single differentiable current-model latent trajectory and adds two kinds of supervision to it: correct-versus-wrong answer readout contrast to decide where stronger supervision is needed, and relevant-versus-mismatched visual evidence contrast to define what those positions should preserve. Both branches are training-only; architecture and inference are unchanged.
-
Demonstrates cross-family, cross-scale applicability. Six backbones spanning Qwen2.5-VL/Qwen3-VL, InternVL3, and Gemma-3, with a five-task average of 63.7% on Qwen2.5-VL-7B, the highest among evaluated latent-reasoning methods there.
-
First scaling of latent visual reasoning to frontier size. Training on Qwen3-VL-235B-A22B (235B total parameters) improves over LVR-SFT on all three evaluated benchmarks — the authors state this is the first demonstration of visual reasoning in latent space trained at frontier scale.
Main Findings
-
ReaLVR leads on Qwen2.5-VL-7B. Five-task average 63.7% ± 0.2 over three random seeds, versus 59.7% for LVR-SFT (+4.0), 60.4% for LVR-RL (+3.3), 60.8% for Monet-SFT (+2.9), 62.2% for Monet-RL (+1.5), and 62.9% for ILVR-Stage2 (+0.8). Per-task: MMVP 72.0, BLINK 55.8, HR-4K 71.8, HR-8K 66.6, MME-RW 52.2. Gains over LVR-SFT cover all five tasks, led by MMVP (+8.4), HR-8K (+3.3), and HR-4K (+2.8).
-
Gains hold across Qwen3-VL sizes. Five-task averages of 61.2% at 8B and 65.2% at 30B (the 30B result is 1.1 points above LVR). At 235B, ReaLVR scores 81.9 on MMVP, 75.4 on BLINK, and 71.0 on MME-RealWorld; HRBench was not evaluated, so no five-task mean is reported there.
-
Gains hold across model families. Five-task means of 58.1 on InternVL3-8B and 41.6 on Gemma-3-12B, exceeding LVR-RL by 3.1 and 2.6 points respectively. Gemma uses a fixed 256-token, single-tile image representation.
-
Vanilla LVR has weak counterfactual sensitivity. Across four edit types with 512 original–edited pairs each and a fixed question, the correct answer changes in 81.45%–86.33% of pairs, but LVR changes its prediction in only 5.66%–13.09%. The gap between ground-truth and LVR prediction-change rates is 68.36–80.67 percentage points. ReaLVR's mean-pooled latent distance between original and edited inputs is 0.13–0.34, versus below 0.0005 for LVR.
-
ReaLVR's latent variation is more position-concentrated. The top-token variation gap is 0.07 for ReaLVR, 0.02 for Monet, and 0.01 for each LVR variant.
-
Latent states carry answer-relevant information without being verbalizable. Applying the Gemma-3-12B vocabulary head to latent vectors from 47 questions with 8 states each (376 projections total) returns "<" as the top-1 token in all cases, with probability 1.0; the latent states align with the
<|lvr_end|>unembedding direction at cosine 0.41 versus 0.02 for a random vocabulary row. Those same states are informative: a linear probe recovers the BLINK task label from the mean latent at 99.9% accuracy. -
Answers depend locally on attributed latent tokens. Replacing the eight most answer-attended latent tokens while holding other states fixed lowers ReaLVR's correct-answer probability from 0.70 to 0.59. In the paper's opening figure, replacing the top-8 tokens ranked by answer-to-token attention reduces correct-answer probability by 4 and 11 percentage points, with the larger drop for ReaLVR.
-
Stronger target-region attention. Target-region enrichment (attention on the annotated region divided by attention on same-area background windows) reaches about 2 in middle layers for ReaLVR, versus near 1.6 for Monet and at or below 1.3 for LVR. Correct and incorrect ReaLVR responses have similar curves, so spatial alignment alone does not explain correctness.
-
Short latent spans suffice. An inference-budget sweep over K = 0 to 20 shows a short span captures much of the benefit, with K = 8 giving the highest mean accuracy.
Methodology in Plain English
The paper starts by auditing what latent tokens actually do. The authors perturb images in ways that should change the answer, then measure whether the latent trajectory and the model's answer move with them — and find they largely do not. A second test looks at attention: latent queries read weakly from the image regions that overlap the annotated region of interest, and answer queries read weakly from latent states.
The proposed fix, ReaLVR, builds on the standard two-stage LVR recipe. Stage 1 already supervises latent states to reconstruct target visual features derived from a region annotation, but under teacher forcing — the model is handed those targets rather than producing them itself. Stage 2 then uses GRPO, which scores only the generated text. ReaLVR's insight is to supervise the trajectory the model actually generates at inference time.
Concretely: for each example, the frozen behavior policy samples a group of completions, which yield GRPO rewards and a set of the model's own wrong answers. Alongside that policy update, ReaLVR regenerates one latent trajectory from the current model, feeding each hidden state back into the next step and keeping gradients through the whole recurrence. Two signals are then computed on this shared trajectory:
-
Where to supervise: the ground-truth answer and each sampled wrong answer are teacher-forced separately after the shared latent span. Attention from their content tokens onto each latent position gives two readouts; the positive difference — correct answer minus mean wrong answer — becomes the per-position credit. Positions read more under the correct answer get more weight; when no valid wrong answer exists, the weight is zero. Weights are detached so the model cannot game the loss by shrinking them.
-
What to preserve: a positive visual prototype is pooled from the image tokens under the region mask (or the whole image if no mask is available), and negatives are pooled the same way from mismatched examples. The loss pushes each latent state to be more similar to the positive prototype than to its nearest negative, by a target margin. The vision encoder and connector stay frozen so the prototypes are fixed targets.
The two are combined into one weighted margin loss added to the standard GRPO objective. Nothing changes at inference: the model generates K latent tokens and decodes the answer exactly as before.
Training details reported: G = 8 completions sampled at temperature 0.6; the paper's main text states training used 800 AMD MI250X GPUs (128 GB each), while the appendix describes each MI250X as a dual-die package exposing two 64 GB GCDs, with the 235B model trained on 200 nodes (800 MI250X, 1,600 GCDs) and smaller backbones on 64 nodes (256 MI250X, 512 GCDs). DeepSpeed ZeRO-3 with no offload, bf16, gradient checkpointing, PyTorch SDPA, sequences packed to 4,096 tokens, 128–5,120 visual tokens per image, AdamW at peak learning rate 1e-5 with cosine schedule, warmup 0.03 and weight decay 0.1, and an MSE reconstruction weight of 0.1 in Stage 1.
Why This Matters
Impact on research. The paper reframes evaluation for latent reasoning: a correct final answer does not prove the latent tokens learned anything useful about the image. It supplies a diagnostic vocabulary (variation, grounding, fixed-context replacement) and shows that answer-level reward alone under-specifies what a latent trajectory should preserve. It also extends latent reasoning from 7–8B models to a 235B backbone, which is the first such demonstration the authors are aware of.
Real-world applications (mapped to the benchmark families the paper evaluates):
- Fine-grained visual discrimination, matching the MMVP evaluation — for example product inspection, quality control, or medical-image triage where subtle cues decide the answer.
- Spatial and relational reasoning, matching BLINK — for robotics, navigation, and manipulation, where "left of" and "behind" determine the action.
- High-resolution perception, matching HRBench-4K/8K — for document, satellite, or aerial-image analysis where evidence occupies a small fraction of the pixels.
- Real-world scene understanding, matching MME-RealWorld — for accessibility tools that describe or answer questions about complex everyday photographs.
Industry relevance. The method is architecture-agnostic, changes nothing at inference, and was validated across three unrelated model families with different vision encoders and language backbones. That makes it a drop-in addition to existing RL post-training pipelines rather than a redesign. The finding that latent states support 99.9%-accurate linear probing while never producing a readable rationale is also directly relevant to deployment: these models are effective but not interpretable by reading their intermediate tokens.
Future Directions
-
Closing the interpretability gap. All 376 inspected latent projections collapse to a single token, yet a linear probe recovers the task label at 99.9%. Understanding what these states encode without a readable rationale remains open, and the authors note the closing-tag prior learned in Stage 1 may be masking the true content.
-
Extending the analysis to full trajectory effects. The fixed-context replacement test isolates local dependence on the eight most answer-attended tokens; the authors note that regenerating later states could introduce further effects that this design does not capture.
-
Separating the two supervision branches more sharply. Appendix J.3 reports a supervision-mass-matched 2×2 test of answer-based position allocation versus positive-versus-negative visual evidence, which is the natural next decomposition of which component does the work.
-
Broadening scale and benchmark coverage. HRBench was not evaluated at 235B, so the frontier-scale claim rests on three tasks. Whether the approach continues to pay off above 235B, and on reasoning tasks beyond discrimination, spatial, and high-resolution perception, is untested in the reported content.
Target Audience
Researchers and engineers working on multimodal LLM post-training, reinforcement learning for language models, and reasoning in continuous latent space. It is also relevant to practitioners who need to know whether a model's correct answer reflects genuine use of the image, and to interpretability researchers interested in what continuous reasoning states encode when they do not decode to language. Readers without a background in RL fine-tuning or vision-language architectures will find the diagnostic sections (counterfactual image edits, fixed-context replacement, target-region enrichment) more accessible than the objective formulation.
Authors’ abstract
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.