Research
Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts Overview Research area: Efficient large language model architecture and training — specifically recurrent ("looped") Transformer d

- arXiv
- 2610.01153
- Published
- 2026-10-01
- Authors
- Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu
AI summary
Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-ExpertsOverview
Research area: Efficient large language model architecture and training — specifically recurrent ("looped") Transformer depth combined with sparse Mixture-of-Experts (MoE) layers.
Technical level: Advanced. The paper assumes familiarity with residual streams, pre-norm Transformers, MoE routing and load balancing, and FLOP accounting for attention and expert layers.
Scope: The paper diagnoses why looped MoE models stop improving after two iterations and proposes LOOM, a training recipe that stabilizes and diversifies recurrence to make loop depth a usable scaling axis from 100M to 1.7B parameters.
What This Paper Is About
Looped Transformers reuse the same block of layers several times, which adds effective depth without adding parameters. When combined with MoE, this idea has stalled in practice: extra loops tend to hurt rather than help, so prior work typically caps recurrence at two loops. This paper asks why deeper looping fails and offers a recipe that lets a 700M model improve up to 5 loops under matched compute, and a 1.7B model improve up to 9 loops when compute is allowed to grow.
Key Contributions
-
A diagnosis of two failure modes. The authors identify the curse of depth (residual updates accumulate across loops, inflating hidden-state variance and destabilizing deep recurrence) and expert selection collapse (shared routers pick nearly the same experts in every loop, so extra iterations add computation without computational diversity).
-
LOOM, a two-part recipe. Built on the principle that each loop should contribute new computation while keeping the recurrent state stable, LOOM stabilizes recurrence with residual scaling and embedding re-injection, and diversifies it with per-loop routers and a Looping Residual.
-
Segmented backpropagation for affordable deep recurrence. The H loops are split into segments of at most K=3 loops, each supervised by a language-modeling loss, with the residual state and global memory detached and passed forward.
-
Empirical evidence that recurrence scales. Across 100M, 350M and 1.7B parameter scales and up to 60B training tokens, LOOM scales to 9–12 loops, and the 1.7B model peaks at 9 loops.
Main Findings
-
Near-iso-FLOP best point is 5 loops. On the M=10, E=80 backbone, all models have 700M parameters and are trained on 10B tokens, with per-token layer cost f(H,k) = H(6+3k) anchored at f₁ = 84 and top-k adjusted to hold cost roughly constant. The non-looped baseline reaches 18.36 validation perplexity and 38.84% average zero-shot accuracy; 2 loops give 17.37 and 39.00%; 4 loops give 16.58 and 39.53%; 5 loops give the best perplexity, 16.54, and 39.53% average accuracy; 6 loops give 16.57 and 39.50%. Four loops already beat the baseline, two-loop, and three-loop settings.
-
Without FLOP matching, deeper keeps winning. At the 1.7B scale trained on 60B tokens, the non-looped baseline reaches 9.62 perplexity and 42.4% average accuracy. LOOM at 3 loops gives 8.94 and 43.9%; 6 loops gives 7.91 and 46.7%; 9 loops is best at 7.77 perplexity and 47.7% average accuracy; 12 loops gives 7.84 and 47.5%. The 9-loop model corresponds to an unrolled scale of roughly 15B parameters (1.7B × 9).
-
Stability holds across scales. In the non-iso-FLOP sweeps, LOOM at ~100M improves from 26.26 baseline perplexity and 36.73% average accuracy to 19.26 and 39.11% at 12 loops; at ~350M it improves from 20.07 and 37.96% to 14.86 and 41.42% at 12 loops. The highest seven-task average in that sweep appears at 12 loops, though gains saturate.
-
Partial fixes are not enough. "No tech" (unmodified looping) becomes unstable with depth, reaching 1311, 1312 and 1587 perplexity at 6, 9 and 12 loops at ~100M (marked with an unrecoverable loss spike). "Res. scale" alone mitigates degradation but still spikes at larger loop counts (33.78 at 9 loops, 35.64 at 12 loops at ~100M). "Embed inject" alone degrades substantially (141.5 at 9 loops and 186.4 at 12 loops at ~100M). All three underperform the non-looped baseline even at shallow recurrence.
-
Each component matters (ablation). Removing residual scaling from the 9-loop ~350M LOOM raises validation perplexity from 18.62 to 24.63; removing embedding re-injection raises it to 26.05; removing the MoE RMSNorm raises it to 26.50; removing the Looping Residual raises it to 19.25; replacing per-loop routers with a shared router ("w/o routing refresh") raises it to 19.83 (remaining entries for that row are not present in the provided text).
-
Loop-specific routers do change expert usage. In a 9-loop ~350M model, loop-specific (independent) routers produce lower cross-loop cosine similarity of expert-load distributions and greater cross-loop variability of per-expert load than shared routers, indicating more diverse expert utilization.
-
Segmented backpropagation cuts memory and time. On the ~350M model with K=3, from H=3 to H=12 segmentation keeps peak memory near 12.6 GiB, while unsegmented training grows from 12.5 to 19.3 GiB. At H=12, reported training time falls from 26.1 to 12.6 hours. At H=9, the unsegmented run degrades severely, with evaluation loss around 4.4 and average accuracy around 34%. Segmented runs also show lower evaluation loss and higher seven-task accuracy.
-
Stabilizing mechanisms are complementary. Residual scaling alone reduces activation variance but yields higher training loss than the full recipe in a 9-loop ~350M model trained on 10B tokens, showing variance control alone is insufficient.
Methodology in Plain English
LOOM takes an M-layer pre-norm MoE decoder and walks through layers 1 to M repeatedly, H times, so effective depth is M×H. Attention and expert weights are shared across loops, but each loop gets its own router (loop-specific sigmoid routers in float32). Adding loops therefore adds router parameters but does not duplicate attention or expert weights.
Two mechanisms keep the recurrent state stable. First, every attention and MoE branch output is multiplied by γ = λ/(H√M), with λ fixed at 0.5. The 1/H factor limits accumulation across loops and the 1/√M factor controls variance across physical layers. Second, starting at loop t = 2, the input embedding is mixed back into the residual stream as h₀ᵗ = (1−g_t)h_Mᵗ⁻¹ + g_t x, where g_t = λ/(t√M), so progressively more weight goes to the recurrent representation while the anchor signal stays well scaled. A further RMS normalization is applied to the MoE output.
Two mechanisms diversify computation. Per-loop routers decouple routing decisions across iterations while expert weights stay shared. The Looping Residual keeps two fixed-decay exponential moving averages (β = 0.5) of scaled attention outputs: a global memory spanning all layers and loops, and a local memory aggregating attention outputs within the current loop. Both are written recursively as N ← βN + o and D ← βD + 1, with the summary r = N/D, so only one accumulator tensor and one scalar are stored per memory regardless of loop count. These summaries replace the direct attention residual update.
Training uses segmented backpropagation: loops are grouped into segments of at most K = 3, each segment receives supervision at its final loop, and gradients are restricted to that segment, with the residual state and global memory detached between segments. This keeps forward information flowing through all H loops while limiting backward depth.
Experiments use a Llama-style pre-norm decoder with grouped-query attention, RoPE (θ = 10⁴), and sparse SwiGLU MoE layers, pretrained on FineWeb-Edu at sequence length 1,024 with a 65,536-token vocabulary. Three scales are used: ~100M parameters (activated ~70M, M=6, hidden size 384, 24 routed experts plus 2 shared, top-k 6 = 4+2, ~5B tokens), ~350M (~141M activated, M=10, hidden size 512, 30 routed plus 2 shared, top-k 8 = 6+2, ~10B tokens), and ~1.7B (~0.63B activated, M=15, hidden size 1280, 30 routed plus 2 shared, top-k 8 = 6+2, ~60B tokens). All scales use per-loop sigmoid routers in float32, z-loss 10⁻³, capacity factor 1.5, load-balance bias learning rate 2×10⁻³, λ = 0.5, β = 0.5, K = 3, peak learning rates of 4×10⁻⁴, 2×10⁻⁴ and 6×10⁻⁵ respectively, AdamW with β₁ = 0.9, β₂ = 0.95, weight decay 0.1, gradient clipping at 1.0, a 10/80/10 warmup–stable–decay schedule with a 0.1× peak learning-rate floor, and a global batch of 1,024.
Evaluation uses seven zero-shot commonsense tasks: PIQA, SIQA, HellaSwag, WinoGrande, ARC-c, ARC-e and OBQA. Length-normalized accuracy (acc_norm) is reported for OBQA, ARC-c, ARC-e, HellaSwag and PIQA, and raw accuracy (acc) for WinoGrande and SIQA. Baselines are the non-looped 1× model and three recurrent variants matched in backbone, data and loop count: "No tech," "Res. scale" and "Embed inject."
Why This Matters
The paper reframes recurrent depth as a genuine, compute-efficient scaling axis for MoE models rather than a curiosity limited to two iterations. If shared expert weights can be re-walked many times, models can gain effective depth without proportionally increasing stored parameters, which shifts the trade-off between parameter count and inference/training compute.
Potential real-world applications (these are implications of the results, not application studies reported in the paper):
- Serving setups where memory capacity, not compute, is the binding constraint, since loops reuse expert and attention weights.
- Training pipelines that want more capability per stored parameter under a fixed accelerator memory budget, aided by the segmented backpropagation memory savings.
- Inference-time or test-time compute scaling, extending recurrence at deployment rather than retraining larger models.
- Research on sparse, routed architectures where the diversity of expert usage across iterations can be controlled explicitly.
Industry relevance: MoE LLMs are the dominant design for frontier-scale training, and anything that improves their compute-to-quality trade-off is directly relevant to pretraining cost and serving hardware. The reported reductions in peak memory (12.6 GiB versus up to 19.3 GiB) and training time (12.6 versus 26.1 hours at H=12 in the ~350M setting) speak to practical infrastructure budgets, and the loop-count hyperparameters (λ, β, K) were held fixed across 100M to 1.7B scales, which matters for teams that do not want to re-tune per scale.
Future Directions
- How far can loops go? The paper reports saturation at 12 loops in the non-iso-FLOP sweep and a best point at 9 loops for the 1.7B model; whether depth beyond this recovers gains, and under what stabilization, is open.
- Do the fixed hyperparameters hold at larger scales? λ = 0.5, β = 0.5 and K = 3 were kept constant from 100M to 1.7B; whether they remain optimal for models and token budgets beyond the ~60B-token, ~1.7B-parameter regime is untested here.
- Can the stabilizing or diversifying mechanisms stack with other approaches? Related work cited (such as mHC's doubly stochastic residual mixing, Attention Residuals, and inter-loop mixing) is not combined with LOOM in these experiments.
- Does routing diversity translate into capability? The paper quantifies lower cross-loop similarity and higher cross-loop variability with loop-specific routers, but does not establish a direct causal link between those routing statistics and downstream task performance.
Target Audience
Researchers and engineers working on LLM pretraining efficiency, sparse MoE architectures, and recurrent or weight-shared Transformer designs. It will be most useful to readers already comfortable with residual-stream analysis, MoE gating and load balancing, and FLOP-matched benchmarking, since the argument rests on interpreting activation-variance plots, expert-load similarity matrices, and iso-FLOP comparisons.
Authors’ abstract
Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available https://github.com/hed-ucas/LOOM.