Research
Scaling Laws for Looped Mixture of Experts
Overview Research area: Machine learning / large language model scaling laws, combining two efficiency axes: looped (recurrent-depth) transformers and Mixture-of-Experts (MoE) sparsity. Technical leve

- arXiv
- 2609.40316
- Published
- 2026-09-30
- Authors
- Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
AI summary
Overview
Research area: Machine learning / large language model scaling laws, combining two efficiency axes: looped (recurrent-depth) transformers and Mixture-of-Experts (MoE) sparsity.
Technical level: Advanced. The paper fits parametric power-law loss models and derives reduced-form limits of those laws; familiarity with Chinchilla-style scaling laws, MoE routing, and weight-tied recurrence is assumed.
Scope: The paper introduces a single scaling law that jointly predicts held-out loss as a function of model size, training data, recurrence (number of loops), and sparsity (expert count), then uses the fitted law to prescribe compute- and memory-optimal architectures and validates the design on downstream benchmarks and at trillion-token training scale.
What This Paper Is About
Existing scaling laws capture either recurrence (looped transformers) or sparsity (Mixture-of-Experts), but not both together. Prior loop laws assume the effective capacity gained from looping grows without bound, while concurrent looped-MoE studies pin recurrence at a fixed value (for example R=2) in their primary scaling analyses. This paper asks how much effective capacity each recurrent pass actually adds, whether that gain saturates, and whether greater MoE sparsity raises and sustains it — then builds a unified law that models all four axes at once.
Key Contributions
- Loop Scaling Laws — a unified law over model size, data, recurrence, and sparsity. The authors state this is the first scaling law to jointly model recurrence and sparsity alongside model size and data, replacing the model-size term N with a recurrence-dependent effective parameter count N_eff(R, m).
- A bounded, sparsity-conditional recurrence mapping. The mapping N_eff(R, m) = N_act + κ1(m)N_loop(1 − e^{−(R−1)/κ2(m)}) treats each recurrent pass as adding fewer effective parameters than the last, converging to a finite asymptote of κ1m^{−θ}N_loop rather than growing without bound. Sparsity, expressed through the active-parameter ratio m = N_act/N_total, raises the asymptote (via κ1(m)) and stretches how many passes the gain is sustained over (via κ2(m)), with κj(m) = κj m^{−θ}.
- Recovery of prior laws as special cases. At R = 1 the law reduces to the standard dense or MoE scaling law; at E = 1 it reduces to the dense loop law; at E = 1 and R = 1 it reduces to the standard dense law (Propositions A.1 and A.2).
- A design recipe plus empirical validation. The fitted law selects compute-optimal recurrence at fixed sparsity, memory-optimal expert count at fixed recurrence, and their joint optimum under both budgets. The authors validate the predictions across 14 downstream benchmarks and at trillion-token training scale.
Main Findings
- The bounded mapping predicts held-out loss better than unbounded alternatives. Under a dense sweep fitted on R ≤ 8 and extrapolated to R = 16, held-out RMSE is 0.0092 for the bounded mapping versus 0.0313 for the power-law mapping and 0.2566 for the linear mapping on held-out R; 0.0128 vs 0.0211 vs 0.1106 on held-out N; and 0.0049 vs 0.0073 vs 0.0882 on held-out D.
- Sparsity conditioning further improves prediction. In the MoE comparison (fitted on R ≤ 8, evaluated on held-out R = 16), the sparsity-conditional mapping N_eff(R, m) achieves held-out RMSE of 0.0100 on R, 0.0047 on E, 0.0050 on N, and 0.0043 on D, lower than the linear (0.2847, 0.0896, 0.0656, 0.0824) and power-law (0.0480, 0.0113, 0.0077, 0.0074) mappings and the sparsity-independent bounded mapping N_eff(R) (0.0190, 0.0075, 0.0065, 0.0058).
- Recurrence adds effective capacity up to a finite, sparsity-dependent asymptote; sparsity raises and sustains it. At fixed active parameter count, the effective-parameter multiplier ρ_eff increases with recurrence but approaches a finite asymptote, and greater sparsity raises that asymptote and delays saturation over more recurrent passes (shown at N_act = 1.0B).
- Higher sparsity improves the IsoFLOP frontier at matched compute, and recurrence pushes it further given more compute. Comparing E = 1 (dense) to E = 8 (MoE), lower loss is attained at the same compute budget. Recurrence R = 2 at 10^21 FLOPs achieves lower loss than R = 1 at 5×10^20 FLOPs for both E = 1 and E = 8.
- Larger training-compute budgets and greater sparsity favor higher recurrence. Compute-optimal R* rises with training-compute budget and is higher for smaller active models; across the F_train–N_act design space, greater sparsity shifts the compute-optimal regime toward higher recurrence.
- Larger weight-memory budgets and higher recurrence favor greater sparsity. Memory-optimal E* increases with memory budget, and higher recurrence favors a larger E* under the same memory.
- At fixed compute, tight memory favors recurrence and more memory favors sparsity. Under a 5×10^21 FLOP budget, a 1.0B active model with E = 8 and R = 2 is selected at 3 GB, whereas a 1.6B active model with E = 16 and R = 1 is selected at 10 GB.
- Sparsity and recurrence give complementary downstream gains. On 14 benchmarks, sparsity at fixed R = 1 consistently improves both Overall and Reasoning across 0.3B, 0.6B, and 1.0B active sizes; at 5×10^20 FLOPs the 0.3B MoE models with E ≥ 8 outperform the 1.0B dense model. Recurrence at fixed E = 8 is more reasoning-oriented: at 10^21 FLOPs the 0.3B model with R ≥ 4 matches or exceeds the 0.6B model with R = 1, and the 0.6B model with R ≥ 3 matches the 1.0B model with R = 1. Sparsity delivers roughly 3× active-parameter efficiency on Overall, and recurrence roughly 2× total-parameter efficiency on Reasoning.
- Gains hold at trillion-token scale with test-time recurrence. At matched training compute (approximately 1.5×10^22 FLOPs), the paper reports a comparison between an A0.6B-2.9B non-looped MoE and an A0.3B-1.3B LoopMoE, with relative inference compute normalized to the non-looped baseline (1.0× for the A0.6B MoE at R = 1, 0.4× for the LoopMoE at R = 1). The abstract states that at matched training compute a looped MoE with law-derived recurrence matches a roughly 2× larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence. The provided text of Table 2 is truncated after the beginning of the second LoopMoE row, so the full per-benchmark results across test-time recurrence values are not available here.
Methodology in Plain English
The authors start from the standard two-term loss formula that predicts loss from model size and training tokens. They then swap the model-size term for an "effective parameter count" that accounts for looping. Unlike earlier approaches that treat every extra loop as adding a fixed chunk of capacity (or a chunk that grows as a power of the loop count), they assume each loop adds less than the one before, so total gain flattens out at a ceiling, much like a curve approaching a wall.
They then make that ceiling depend on how sparse the model is. In a dense looped model, every pass reuses exactly the same weights. In a looped MoE, the router can send a token to different experts on different passes, so each loop can reach a different slice of the parameters. Sparsity is expressed as the ratio of active to total parameters, and that ratio scales both the height of the ceiling and how quickly it is reached.
To fit and test the laws, they train models across a grid of four axes at once: active parameters (0.3B, 0.6B, 1.0B), training tokens (100B to 500B), expert count (1, 2, 4, 8, 16), and recurrence (1, 2, 3, 4, 6, 8), using middle-cycle looping that leaves the first two and last two layers unshared. They fit on runs up to R = 8 and check predictions on held-out R = 16, unseen model sizes, unseen data amounts, and unseen expert counts, scoring with RMSE.
Finally, they use the fitted law as an optimizer: given a training-compute budget and a weight-memory budget, they search over configurations, apply constraints (including a tolerance ε set to the fitted-law RMSE, so that a further loop or more experts is only counted as an improvement if the predicted loss drop exceeds fitting error), and pick the lowest-predicted-loss design. These predictions are then checked against actual downstream evaluations on 14 benchmarks spanning reasoning, science, commonsense, reading, and knowledge.
Why This Matters
Impact on research. The paper reframes recurrence and sparsity as a single joint scaling problem rather than two independent ones, and it argues against the unbounded-gain assumption used in earlier loop scaling laws. It provides reduced forms that recover Chinchilla-style dense laws, MoE laws, and dense loop laws, offering one functional form that prior work can be mapped onto. It also supplies a concrete recipe for choosing architecture under resource constraints, which turns a predictive law into a design tool.
Real-world applications:
- On-device and edge deployment, where weight memory is the binding constraint: the paper's memory-optimal analysis maps memory budgets directly to expert count and recurrence choices.
- Cloud-served frontier models, where sparsity at matched compute can substitute for a larger dense model.
- Test-time compute control, since recurrence can be varied at inference to trade additional compute for quality on demand without changing stored parameters.
- Budget-constrained training planning, where a compute budget and memory budget are given and the law selects active size, expert count, and recurrence.
Industry relevance. The work is authored at Meta AI and explicitly situates itself relative to frontier and edge MoE deployments, including Mixtral, DeepSeek-V3, Qwen3MoE, Gemini, MobileMoE, and Apple's AFM 3 Core Advanced. Its central claim — trading extra inference compute for parameter efficiency at fixed quality — is directly relevant to serving economics, memory footprint, and the cost balance between training and inference.
Future Directions
- Other looping strategies. The authors use standard middle-cycle looping and state explicitly that their formulation is agnostic to the specific looping strategy, leaving other strategies to future work.
- Deployment memory beyond weights. The memory-optimality analysis considers weight memory only (M_weight = b_w N_total / 8); KV cache and other runtime state are stated to be outside the paper's scope.
- Validation of the asymptotic limits. The finite effective-parameter bound and the extreme-sparsity linear limit are described as properties of the mapping rather than guarantees beyond the observed recurrence range, leaving open whether they hold at larger R and larger expert counts.
- Broader integration of further scaling axes and configurations. The joint-optimum search is run over a model ladder spanning active model size, expert count, and recurrence; extending the framework to additional design axes is an open question the paper does not resolve.
Target Audience
Researchers and engineers working on scaling laws, efficient transformer architectures, or MoE systems who need to decide how to allocate a fixed compute and memory budget across model size, data, recurrence, and expert count. It is most useful to readers comfortable with fitted loss curves and MoE routing terminology; readers looking for an introductory treatment of scaling laws will find the derivations and RMSE tables demanding. Practitioners focused on memory-constrained deployment (edge or on-device) and on inference-cost tradeoffs will find the design-recipe sections directly applicable.
Authors’ abstract
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.