Skip to content
AI.info

Research

Decoding Looped Transformers Better for (Almost) Free

Overview Research area: Efficient inference and decoding for Looped Transformers (a recurrent-depth language model architecture), specifically training-free contrastive decoding. Technical level: Inte

Decoding Looped Transformers Better for (Almost) Free
arXiv
2610.02185
Published
2026-10-01
Authors
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

AI summary

Overview

Research area: Efficient inference and decoding for Looped Transformers (a recurrent-depth language model architecture), specifically training-free contrastive decoding.

Technical level: Intermediate. The paper assumes familiarity with transformer decoding, logits, hidden states, and contrastive decoding; the core method itself is a simple linear extrapolation.

Scope: The paper introduces LoopCD, a training-free decoding framework that uses an early recurrent pass of a looped Transformer as a weak reference to guide the final pass, improving accuracy and enabling roughly half the recurrent iterations at equal or better accuracy.

What This Paper Is About

Looped Transformers reuse one shared block of layers repeatedly, so each recurrent pass produces an intermediate hidden state that could be decoded into a next-token prediction — but standard decoding throws all of those away and only uses the final pass. The paper asks whether those discarded intermediate states can be used as a "weak" reference to guide the "strong" final prediction, in the same spirit as contrastive decoding, without any extra training or auxiliary models. The goal is better decoding quality from the same checkpoint, and ultimately fewer recurrent iterations for the same or better accuracy.

Key Contributions

  1. LoopCD, a training-free contrastive decoding method. It exploits the observation that a looped Transformer inherently produces aligned weak-and-strong predictions across recurrent depth, requiring no auxiliary model, perturbed prompt, or external training.
  2. Two guidance spaces. LoopCD-Logits contrasts logits from the first and final states (one extra output pass), while LoopCD-Hidden combines the two hidden states before the coda layers and language modeling head, eliminating extra output projection overhead entirely (zero output overhead).
  3. Demonstrated gains across four model families. Ouro, Huginn, Parcae, and Looped-Qwen3 all improve on mathematical reasoning, code generation, and multiple-choice benchmarks at full recurrent depth.
  4. Compute reduction. The gains enable halving recurrent iterations while matching or exceeding full-depth unguided baselines, eliminating 22.5% to 48.2% of forward FLOPs.

Main Findings

  • Large reasoning gains at full depth. Adaptive LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, AIME 2025 pass@1 from 49.58% to 56.88%, and OlympiadBench pass@1 from 64.05% to 67.29%. Across both Ouro-Thinking models, fixed and adaptive rules improve all evaluated benchmarks by 5.27 to 7.33 points on pass@1 and 2.53 to 4.12 points on pass@10.

  • Code generation improvements, with the hidden variant strongest. For Huginn-0125 at R=32, LoopCD-Hidden raises HumanEval pass@1 from 22.56% to 31.71%, ahead of adaptive LoopCD-Logits (23.17% to 28.66%) and fixed (27.44%). At R=16, LoopCD-Hidden moves HumanEval pass@1 from 21.95% to 28.05%, matching the adaptive logit score.

  • Consistent multiple-choice gains across architectures. Every evaluated model improves its seven-benchmark mean under LoopCD-Logits, with the adaptive rule leading across nearly all configurations. LoopCD-Hidden raises the mean by +0.83 and +0.75 points on Huginn at R=32 and R=16 (exceeding fixed logit gains of +0.74 and +0.62) and by +0.85 points on Parcae-1.3B.

  • Robustness to different recurrent training objectives. Ouro supervises readouts after every pass, Huginn and Parcae sample recurrent depths during pre-training, and Looped-Qwen3 retrofits recurrence onto a frozen non-recurrent model; LoopCD helps in all four regimes.

  • Half the iterations, equal or better accuracy. Halving recurrent iterations costs the unguided model 0.17 to 1.29 points on the seven-benchmark multiple-choice mean; applying LoopCD at half depth adds 0.75 to 1.29 points, closing the gap in all six evaluated settings. Huginn-0125 at sixteen of its thirty-two iterations beats its full-depth unguided baseline by 1.02 points under LoopCD-Logits and 0.51 under LoopCD-Hidden; Looped-Qwen3 matches its full-depth baseline at half iterations under LoopCD-Hidden (+0.00 points).

  • Measured compute savings. The guided model at halved depth requires only 0.52 to 0.78 of the original forward FLOPs, accounting for all guidance overhead. Across reduced-depth settings, LoopCD eliminates 22.5% to 48.2% of total forward FLOPs.

  • The weak-to-strong premise is validated empirically. The prediction after the first iteration trails the final converged prediction by 4.0 to 23.7 points on the seven-benchmark mean, with accuracy climbing across iterations. Divergence from the final prediction falls from 0.71 bits after Ouro-1.4B's first pass to 0.013 after its third, and from 3.0 nats after Huginn's first step to 0.04 after its sixteenth.

  • Gains track disagreement with the final prediction. Huginn's first step picks a different option from the final prediction on 51% of ARC-Challenge questions and yields 3.07 points at ω=0.5; its sixteenth step disagrees on 7% and yields 0.17. Gains drop from 3.07 to 0.17 points as reference depth increases for Huginn and from 3.50 to 1.02 for Parcae-1.3B.

  • The first recurrent step is the best logit reference, but Huginn's hidden reference needs burn-in. For LoopCD-Hidden, Huginn initializes from Gaussian noise, so early hidden states yield -0.58 points before peaking at the sixth step with +0.83 points. Parcae and Looped-Qwen3 start from deterministic representations and effectively use their first state in both spaces.

  • Re-ranking, not temperature scaling, drives the gain. Projecting the logit contrast onto directions parallel and orthogonal to the final logits separates a uniform rescaling from a pure re-ranking. On HellaSwag, the orthogonal component alone yields +1.72 points versus +0.76 for the full update on Ouro-1.4B, and +2.27 versus +1.56 on Ouro-2.6B at ω=0.5.

  • Guidance acts where the model is uncertain. At ω=0.5 on ARC-Challenge, gains are between +6.4 and +13.3 points on the least confident fifth of questions across four models and at most +0.4 on the most confident fifth. Only 5.8% to 14.9% of answers flip, and the net gain of 2.1 to 3.5 points comes strictly from resolving close decisions.

  • Different strength regimes for scoring and generation. Multiple-choice scoring tolerates a wide strength band peaking near ω=0.5, while autoregressive generation needs a narrower window of ω in [0.2, 0.3] because early token shifts compound. Negative ω universally degrades accuracy.

  • Adaptive margin gating widens the usable range. Scaling strength by the top-two probability margin (ω = ω_max[1 − (p_R,(1) − p_R,(2))]) keeps gains positive across ω_max in [0.5, 1.0], where fixed strength overshoots.

Methodology in Plain English

A looped Transformer runs the same block of layers over and over, refining a hidden state before producing a token. The researchers take the state from the very first pass (the weak prediction) and the state after the last pass (the strong prediction), then push the final prediction further in the direction the recurrence was already moving: they extrapolate away from the early state, h' = h_R + ω(h_R − h_1).

For LoopCD-Logits, both states are pushed through the coda layers and the language modeling head, and the same extrapolation happens on the resulting logits. This costs one extra pass through the output layers. For LoopCD-Hidden, the two hidden states are combined before those layers, so the coda and head run only once and there is no extra output cost. Both forms then decode from the softmax of the guided scores.

The strength ω is either a fixed constant per evaluation setting or adaptive: it scales with how close the top two token probabilities are, applying maximum strength only to genuinely contested tokens. Because obtaining the unguided probabilities needed for that rule would require an extra pass, adaptive strength is used with LoopCD-Logits, and fixed strength with LoopCD-Hidden.

Evaluation is done as paired comparisons: guided and unguided runs share the same checkpoint, prompts, shot count, generation limit, stopping rule, answer extractor, evaluator, and recurrent depth. Tables report percentage-point changes from the matched baseline.

Why This Matters

The work shows that a looped Transformer's discarded intermediate computation is a free source of guidance signal, turning one trained network into a built-in weak-to-strong pair. It also demonstrates that a purely inference-time method can substitute for recurrent computation, shifting the accuracy-compute frontier rather than just improving accuracy at fixed cost.

Real-world applications (implications drawn from the paper's evaluated domains):

  • Code assistants: LoopCD improved HumanEval and MBPP execution pass rates, which matters for tooling that generates or completes code under latency constraints.
  • Mathematical and scientific reasoning: Gains on AIME 2024, AIME 2025, and OlympiadBench point to stronger step-by-step problem solving without retraining.
  • On-device or cost-sensitive inference: Halving recurrent iterations while matching full-depth accuracy cuts forward FLOPs by 22.5% to 48.2%, which is directly relevant to serving cost and latency.
  • Deployment on existing checkpoints: Because the method is training-free and needs at most one extra output pass, it can be applied to already-trained looped models without fine-tuning.

Industry relevance: The efficiency claims target inference cost directly, and the method applies across four architecturally distinct model families, including one (Looped-Qwen3) built by retrofitting recurrence onto a frozen Qwen3-4B. Both variants require either one extra output pass or none, which matters for throughput-oriented serving.

Future Directions

  • Making the weak-strong pair more useful through training. The authors note that models trained with their intermediate predictions in mind, as Ouro is, may make the pair more useful still — suggesting recurrent training objectives could be redesigned around guidance.
  • Understanding when early hidden states are usable. Huginn's Gaussian-noise initialization forces a burn-in before hidden-state guidance works, while Parcae and Looped-Qwen3 can use their first state; how initialization choices affect hidden-space guidance remains open.
  • Extending to deeper codas and other output interfaces. The paper's Appendix D.3 analyzes how much of the logit gain the hidden form retains behind a deeper coda, and notes the probability-space identity cannot be applied directly to LoopCD-Hidden because coda layers are nonlinear.
  • Choosing strength automatically. The results show a roughly two-fold difference in tolerable fixed strength between scoring and generation, and adaptive gating widens but does not remove the tuning requirement; a principled per-token or per-task strength schedule is a natural next step.

Target Audience

Researchers and engineers working on efficient inference, recurrent-depth and weight-tied language models, and contrastive or guided decoding. It is also useful for practitioners deploying looped Transformer checkpoints who want accuracy or compute improvements without retraining, and for anyone studying weak-to-strong supervision or inference-time scaling. Readers need working knowledge of transformer decoding, logits, and softmax to follow the method and the ablations.

Authors’ abstract

Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.

Read the original paper