Skip to content
AI.info

Research

LoopVL: Recurrent Visual Intelligence

LoopVL: Recurrent Visual Intelligence Overview Research area: Computer vision and multimodal (vision–language) modeling, specifically recurrent / looped Transformer architectures that reuse shared par

LoopVL: Recurrent Visual Intelligence
arXiv
2609.38426
Published
2026-09-29
Authors
Zhe Qian, Ziyang Gong, Zhongxing Xu, Hehan Li, Zhonghua Wang, Fei Luo, Mingxuan Wang, Xue Yang, Shiwei liu, Yanbiao Ma, Junchi Yan, Jungong Han

AI summary

LoopVL: Recurrent Visual Intelligence

Overview

Research area: Computer vision and multimodal (vision–language) modeling, specifically recurrent / looped Transformer architectures that reuse shared parameters to increase computational depth without increasing parameter count.

Technical level: Advanced. The paper assumes familiarity with Transformer internals (attention, RoPE, PrefixLM masking, SwiGLU), recurrent-depth models, and vision–language training pipelines (alignment, mid-training, supervised fine-tuning, GRPO reinforcement learning).

Scope: The paper asks whether Loop Transformers can be extended from language to vision–language models, and reports one 1B-scale recurrent model (LoopVL-1B) trained from scratch across language pre-training, multimodal training, and post-training, together with an analysis of how visual attention and visual states evolve across recurrent cycles.

What This Paper Is About

Recurrent Transformers increase effective computational depth by applying the same shared modules repeatedly instead of stacking independently parameterized layers, but prior work on this idea has focused on language models. Extending it to vision–language models raises a new issue: after one round of vision–language interaction, both the visual representations and the relative importance of visual tokens can change, so later recurrent steps operate on an evolving multimodal state rather than a static one. The paper builds LoopVL to test this directly — training a recurrent vision–language model from scratch and measuring both its benchmark performance and how its visual processing changes across loops.

Key Contributions

  1. Recurrent vision–language modeling. The authors introduce LoopVL, which combines Module-Loop and Model-Loop computation to iteratively update a unified vision–language state through shared modules, and demonstrate that repeated computation with shared parameters can be applied to multimodal modeling.

  2. Comprehensive validation of multimodal recurrence. They perform vision–language pretraining and post-training of LoopVL and evaluate it across a range of multimodal tasks and recurrent configurations (H1L1, H1L3, H2L1, H2L3 trained independently; H1L1 through H4L3 at inference), providing systematic empirical evidence for extending recurrent Transformers to vision–language models.

  3. Analysis of recurrent visual dynamics. They study how visual information evolves throughout recurrent computation, reporting that later recurrent steps do not merely repeat earlier computation but continuously reorganize and reuse visual evidence. This includes the "Visual Aha Moment," cross-loop token relevance, and visual-attention allocation analyses.

Main Findings

  • Recurrence beats parameter-matched depth. Under the same 0.14T-token training budget, LoopVL (4 recursions, H2L3, 2.47 × 10²¹ FLOPs) improves over Transformer-VL 1B (32 layers, hidden size 1536, 0.92 × 10²¹ FLOPs) on every reported benchmark: MMStar 55.33 → 63.47, RealWorldQA 55.29 → 70.98, ChartQA 51.12 → 74.52, VMCBench 55.10 → 70.90, AI2D 60.01 → 75.49, MathVision 30.59 → 38.49, VisuLogic 21.20 → 27.00, MMK12 40.30 → 49.65. Both models contain the same number of unique Transformer layers, but LoopVL raises executed depth from 32 to 128 layer calls.

  • Competitive with much larger dense baselines. LoopVL approaches or surpasses the larger Transformer-VL 4B variants (Deep: 78 layers, hidden 1792, 2.89 × 10²¹ FLOPs; Wide: 32 layers, hidden 2816, 2.99 × 10²¹ FLOPs) on several benchmarks while using fewer estimated training FLOPs. It exceeds both on MMStar (63.47 vs 61.33 and 60.47) and VisuLogic (27.00 vs 26.60 and 26.30), and trails them on AI2D (75.49 vs 76.13 and 75.65) and VMCBench (70.90 vs 71.20 and 72.00).

  • Training-time recurrence organization matters more than raw depth. In the independent training ablation, H2L3 (128 unrolled layers) scores highest on all five benchmarks. H1L3 and H2L1 both execute 64 Transformer-layer calls, yet H2L1 is better on all five (MMStar 60.80 vs 58.13; RealWorldQA 64.58 vs 59.61; VMCBench 63.80 vs 60.50; AI2D 68.26 vs 66.06; ChartQA 65.12 vs 60.40).

  • Inference must match the training schedule. On the fixed H2L3-trained model, training-matched H2L3 (128 unrolled layers) is best on all five benchmarks. H1L1–H1L3 produce largely unparseable responses and near-zero scores (MMStar 0.47, 0.07, 0.07). With two model cycles, increasing L from one to three raises MMStar from 29.40 (H2L1) to 51.33 (H2L2) to 63.47 (H2L3). Adding cycles beyond training does not help: H3L3 (192 layers) and H4L3 (256 layers) drop MMStar to 58.20 and 24.00.

  • Equal unrolled depth does not imply equivalent behavior. H1L3 and H2L1 both execute 64 layer calls, H2L2 and H3L1 both execute 96, and H2L3 and H4L1 both execute 128 — yet performance differs substantially. For the 128-layer pair, MMStar is 63.47 under H2L3 but 24.67 under H4L1.

  • Visual Aha Moments. Using a 32-sample diagnostic set, the authors measure normalized spatial entropy and Gini concentration of visual attention at every executed layer. Entropy is generally lower during the second cycle but fluctuates across module calls rather than decreasing monotonically. The mean Gini curve rises from 0.405 to 0.858 across the model-cycle transition, which lies between zero-based indices 63 and 64. The authors call this pronounced cross-cycle shift in visual evidence allocation a Visual Aha Moment. They note that a high Gini value describes concentration and does not by itself identify attention sinks or causally important visual tokens.

  • The phenomenon is robust to backpropagation route and to the visual anchor–gate mechanism. A model pretrained from scratch with gradients propagated through all recurrent invocations reproduces the cross-cycle attention shift and boundary concentration, and a second model pretrained from scratch with the visual anchor and query-conditioned visual gate removed shows the same pattern.

  • Visual-state updates help. Four conditions are compared on LogicVista, RealWorldQA, VMCBench-DEV, and MMStar: Normal, two inference-only controls (Block L visual updates, Freeze visual states), and a separately trained read-only variant. Normal achieves the highest accuracy on all four; the retrained read-only variant remains below it. The numeric scores for this comparison are not reported in the available text.

  • Cross-loop reallocation of evidence. Across eight selected examples, the paper aligns the same final H-module layer at the two cycle endpoints (zero-based indices 63 and 127) and marks tokens promoted from the first-cycle bottom half to the second-cycle top quartile, distinguishing tokens inside versus outside the outlined target region.

  • Compact model, competitive results at low token cost. In the 16-benchmark-family comparison, LoopVL-1B is trained on 0.14T tokens while comparison models range from 3T to 36.38T tokens. LoopVL-1B scores 45.19 on LogicVista, 38.67 on MMMU-Pro, 13.14 on BabyVision, 34.76 on VisualPuzzles, 44.71 on HallusionBench, 90.71 on POPE (87.60 POPE-A, 91.92 POPE-P, 92.78 POPE-R), 61.60 on MMK12-Math, 38.49 on MathVision, 20.39 on MathVision-WildPhoto, and 49.65 on MMK12.

Methodology in Plain English

The team built a recurrent vision–language model from a frozen Penguin Vision Encoder, a trainable visual interface, and their own recurrent language backbone pretrained on text using HRM-Text's open-source framework. The encoder produces a 1024-dimensional feature per retained image patch; a two-layer GELU projector maps each feature from 1024 to 1536 dimensions through a 1536-dimensional hidden layer, and the projected embeddings replace reserved visual slots in a sequence containing a conditioning header, image, instruction, and response.

The language backbone holds two separate 16-layer Transformer stacks, L and H, that do not share parameters with each other but each reuse their own parameters across repeated invocations. The default H2L3 schedule runs L → L → L → H, then L → L → L → H again, so a single forward pass performs 128 Transformer-layer calls from only two stored stacks. The H state is initialized from scaled input embeddings and the L state from zero; within each cycle L is updated repeatedly under the current high-level condition, then H incorporates the refined L state. At each new model cycle, a fixed anchor from the initial projected visual embeddings is re-injected into the visual positions of the H state, controlled by a query-conditioned token-wise visual gate and a learnable cycle-specific scale.

Training uses selective backpropagation: the backward graph starts with the final two invocations (L6 and H2) and expands to the final five (L4, L5, L6, H1, H2), leaving L1, L2, and L3 out of gradient computation while all invocations still run in the forward pass.

Training proceeded in stages. Text pretraining used 75B tokens (60B from the HRM-Text base recipe plus 15B additional curated data). Stage 1 (Visual Language Alignment) used LLaVA-559K at 0.33B tokens, a 2048 maximum sequence length, 256–1024 visual tokens per image, and trained only the visual interface with the encoder and language backbone frozen. Stage 2 (Multimodal Mid-Training) used OneVision-1.5 Mid-Training-85M at 49B tokens, a 3072 maximum sequence length, global batch size 128, and jointly updated the language backbone and visual interface. Both stages used AdamW with (β1, β2) = (0.9, 0.95), gradient-norm clipping at 1.0, 3% learning-rate warmup, and zero weight decay. Post-training added supervised fine-tuning on 16.81B tokens (OneVision-1.5 Instruct-22M, OneVision Spatial-4M, LLaVA v1.5 Mix665K, and Vision-R1-Cold, at roughly 77.3%, 17.8%, 3.3%, and 1.5%) and visual reinforcement learning on 0.006B tokens from Vision-R1-RL using GRPO with an answer-correctness reward and a progressive reasoning-length schedule from 2048 to 4096 tokens.

Evaluation covered 16 benchmark families using greedy generation, with a 32-token maximum for short-answer tasks, and used VLMEvalKit plus official benchmark code. To explain the results, the authors compared trained recurrence configurations, swept inference-time schedules on the fixed H2L3 model, and measured visual attention with normalized spatial entropy and Gini concentration at each executed layer.

Why This Matters

Impact on research. The paper provides direct empirical evidence that recurrent computation transfers from language to vision–language modeling, and it introduces a way to interpret recurrent models by exploiting the spatial correspondence between visual tokens and image patches. Because the same physical layer can be compared across recurrent cycles, the authors can distinguish "repeating the same computation" from "applying shared parameters to continuously evolving visual states." The finding that H/L schedule order matters more than unrolled depth — and that inference schedules must match training schedules — is a useful constraint for anyone designing recurrent models.

Real-world applications (as directions this work could inform):

  • Visual reasoning assistants that need to work through a question by progressively re-examining different regions of an image rather than reading it once.
  • Document, chart, and text-rich image understanding, where the model must move attention between local text and global layout.
  • Deployment on constrained hardware, since recurrence yields deeper executed computation without storing proportionally more parameters.
  • Multimodal models trained under tight data budgets, since LoopVL-1B is trained on 0.14T tokens while many compared models use 3T to 36.38T tokens.

Industry relevance. The efficiency angle is the main draw: LoopVL reaches competitive multimodal scores with a roughly 1B-scale language backbone and fewer estimated training FLOPs (2.47 × 10²¹) than the 4B dense baselines it is compared against (2.89 × 10²¹ and 2.99 × 10²¹), and with far fewer training tokens than most of the models in the comparison table. The paper also shows that the gains depend on careful schedule matching between training and inference, which is a practical deployment constraint.

Future Directions

  • Determining why schedule organization matters. H1L3 and H2L1 execute identical numbers of Transformer-layer calls but perform differently, and H2L3 and H4L1 both execute 128 layer calls with MMStar of 63.47 versus 24.67. The paper does not explain the mechanism behind this, and notes that models trained with longer schedules may behave differently.

  • Explaining the degradation from extra cycles. H3L3 and H4L3 execute more computation than the trained H2L3 schedule yet score lower on all five benchmarks, which raises the open question of what limits useful recurrence depth in this checkpoint.

  • Causal validation of the Visual Aha Moment. The authors explicitly state that a high Gini value describes concentration and does not by itself identify attention sinks or causally important visual tokens, leaving the causal role of the reallocated regions unestablished.

  • Extending the visual-state intervention study. The read-only trained variant scored below Normal on all four evaluated benchmarks, but the available text does not report the numeric results; a fuller account of how restricting visual-state updates in later cycles affects different task families would extend this line.

Target Audience

Researchers and graduate students working on recurrent or looped Transformer architectures, vision–language models, and multimodal training pipelines

Authors’ abstract

We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.

Read the original paper