Research
One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
Overview Research area: Computer vision — efficient Vision Transformer (ViT) architectures, recurrent/weight-shared networks, and mixture-of-experts (MoE) parameter composition. Technical level: Advan
- arXiv
- 2610.12448
- Published
- 2026-10-08
- Authors
- Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos
AI summary
Overview
- Research area: Computer vision — efficient Vision Transformer (ViT) architectures, recurrent/weight-shared networks, and mixture-of-experts (MoE) parameter composition.
- Technical level: Advanced (assumes familiarity with ViT blocks, FFN/MoE routing, knowledge distillation, and representation-similarity metrics).
- Scope: The paper introduces reViT, a ViT in which one Transformer block is applied recurrently and its FFN at each depth is produced by mixing a shared bank of experts according to a continuous normalized-depth coordinate, evaluated under supervised ImageNet-1k training and DINOv2 distillation.
What This Paper Is About
Standard ViTs process tokens through a stack of L blocks, each with its own parameters, so parameter storage grows with depth. Prior work (Raptor) showed that a few distinct blocks applied recurrently recover most of a full-depth model's accuracy, but left open the extreme case where a single block is reused at every depth. reViT addresses that case: it restores depth-specific computation not by storing more blocks, but by using normalized recurrent depth to softly merge a shared bank of FFN experts into one dense FFN per step.
Key Contributions
- A recurrent ViT with a continuous depth program. reViT replaces a ViT's depth-wise stack with one recurrent Transformer module (shared attention, LayerNorms, router, and expert bank). A normalized depth coordinate s_t = t/(L−1) is mapped by a small MLP router to soft coefficients over an E-expert FFN bank, and the experts are merged in weight space into one dense FFN applied to every token.
- A controlled comparison of expert mechanisms under a common recurrent setting. Using the same recurrent backbone, task loss, and base recipe, the paper compares seven alternative MoE mechanisms at the nominal budget of one dense FFN per block (U = 1), testing token-dispatch and output-mixture adaptations against weight-space merging.
- Resampleable depth and fixed-depth export. One checkpoint supports multiple tested inference depths by resampling the normalized coordinate interval (elastic-depth training/inference), and for fixed-depth deployment the recurrent block can be materialized as a conventional dense graph of L merged FFNs, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
- Analysis of the learned depth program. The paper characterizes how the router allocates experts over depth, whether the learned gate-expert assignment matters (via gate-column permutations and expert ablations), and how fully the expert bank is used at different model scales (via effective rank).
Main Findings
- Supervised ImageNet-1k: At near-matched inference FLOPs, the E = 4 reViT models score 0.1–0.2 points above four-block Raptor across all scales. reViT-B/16 reaches 83.0% versus DeiT III's 82.8%, and reViT-L/16 reaches 83.8% versus 84.1%, with roughly 7× fewer parameters. The S/16 gap is larger (78.2% versus 79.9%) but narrows with additional experts.
- Parameter reporting: The abstract states reViT-B/16 attains DeiT III accuracy with about 70% fewer stored parameters; the introduction states 73% fewer parameters. Both figures appear in the paper as written.
- DINOv2 distillation: An E = 8 reViT-B/14 student reaches 83.9% ImageNet linear-probe top-1, 45.8 ADE20k mIoU, 0.554 NYUv2 linear RMSE, and 0.483 NYUv2 MLP RMSE, compared with the DINOv2-B/14 teacher's 84.5%, 47.5, 0.567, and 0.486. The E = 1 model reaches 79.4%.
- Distillation beats Raptor on all four probes at E = k ∈ {2,3,4}, for example E = 4 at 83.5% IN-1k / 44.6 ADE20k / 0.552 / 0.491 versus Raptor k = 4 at 83.2% / 43.6 / 0.564 / 0.495 — though the paper notes these are reference comparisons, not controlled ablations, because the distillation objectives differ (Raptor uses intermediate teacher features; reViT uses only final-layer features).
- MoE mechanism comparison, supervised (U = 1): reViT reaches 78.2%, Lory 77.9%, and SMEAR 77.7%, while every other method falls between 67.8% and 70.6%. Higher-budget configurations narrow the gap, with MoEUT reaching 77.6% at U = 4.
- MoE mechanism comparison, distillation: At U = 1, reViT has the best reported ImageNet, ADE20k, and linear-depth results, while SMEAR achieves the lower MLP-head RMSE (0.485 versus 0.491). The Qwen3 run at U = 1 did not converge to a usable solution. At U = 4, ReMoE leads on all four tasks.
- Elastic-depth inference: At L = 12, the depth-only gate reaches 44.4 ADE20k mIoU, 83.4% ImageNet top-1, and 0.493 NYUv2 RMSE, versus approximately 35.6 ADE20k mIoU for a SMEAR-style feature-only gate. ADE20k rises from 40.3 mIoU at L = 8 to 44.9 at L = 16 and stays at 44.7 at L = 24; NYUv2 RMSE improves from **0.
Authors’ abstract
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70\% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.