Skip to content
AI.info

Research

DLoop: Looped Speculative Decoding

Overview Research area: Natural language processing, specifically efficient inference for large language models — the subfield of speculative decoding. Technical level: Advanced. The paper assumes fam

DLoop: Looped Speculative Decoding
arXiv
2610.07659
Published
2026-10-06
Authors
Geonmo Gu, Byeongho Heo, HeeJae Jun, Yoohoon Kang, Sangmin Lee, Sangdoo Yun, Dongyoon Han

AI summary

Overview

Research area: Natural language processing, specifically efficient inference for large language models — the subfield of speculative decoding.

Technical level: Advanced. The paper assumes familiarity with autoregressive generation, draft/target model pairs, verification passes, parallel draft models, and hidden-state reuse in speculative decoding frameworks.

Scope: The paper introduces DLoop, a modification to speculative decoding that performs several drafting stages before verification, and reports wall-clock speedup gains across several existing speculative decoding methods without changing the output distribution.

What This Paper Is About

Speculative decoding speeds up language model generation by having a small draft model propose tokens that a larger target model then checks. The authors observe that when draft models are strong, the target model often accepts every token from a drafting round — yet a verification pass still runs each time, wasting target-model computation that could have gone toward drafting more tokens first.

The goal of DLoop is to let drafting continue across multiple stages while the draft model stays confident, and to verify all the accumulated tokens in a single target-model pass. This is framed as a way to trade more cheap draft-model passes for fewer expensive target-model passes.

Key Contributions

  1. A looped drafting scheme. DLoop restructures speculative decoding so that multiple drafting stages can occur back-to-back before any verification, rather than one drafting stage per verification.

  2. A confidence-based stopping rule. Drafting continues adaptively while the draft model remains confident, and all tokens accumulated across those loops are verified together in one target-model pass.

  3. Loop-aware training. The draft model is trained on its own hidden states for draft tokens that have not yet been verified, which keeps it reliable in the extra drafting stages that DLoop introduces.

  4. A solution for the parallel draft model bottleneck. The paper addresses the specific problem that parallel draft models need target-model hidden states for unverified draft tokens in order to keep drafting, which ordinary adaptive draft-length methods (designed for autoregressive draft models) do not handle.

Main Findings

  • Verifications are often unnecessary. With increasingly capable draft models, the target model frequently accepts all tokens produced in a single drafting stage, meaning the verification pass that follows was not needed to reject anything.

  • Adaptive draft length has a limited scope. Prior methods that decide during decoding how many draft tokens to produce before verification raise speedup only for autoregressive draft models, not for parallel ones.

  • Parallel drafting has a hidden dependency. For a parallel draft model to draft further, it requires target-model hidden states for draft tokens that have not yet been verified — a constraint DLoop is built to work around.

  • Reported speedup range. Across diverse speculative decoding methods — EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules — DLoop improves wall-clock speedup by 5 to 41 percent.

  • Lossless decoding is preserved. The paper states the speedup comes without changing the decoding output, i.e., the target model's distribution is maintained.

  • The trade is deliberate. DLoop spends additional draft-model forward passes in order to reduce the number of target-model forward passes needed for verification.

Methodology in Plain English

Standard speculative decoding alternates: draft a few tokens, then verify them with the big model. The authors noticed that when the draft model is good, the verify step usually comes back with everything approved — so the big model was invoked for nothing, and drafting could have simply kept going.

DLoop changes the loop structure. Instead of drafting once per verification, the draft model keeps drafting in successive stages as long as it feels confident about what it is producing. Only when that confidence drops does the system hand the whole batch of accumulated tokens to the target model for a single verification pass.

Making this work required two pieces. First, the system needs a way to decide when to stop looping and verify — the confidence signal serves that role. Second, a parallel draft model normally cannot keep drafting without the target model's hidden states for tokens that have not been verified yet, since those states are what verification would produce. The authors handle this with loop-aware training, which feeds the draft model its own hidden states for unverified tokens so it stays dependable during the additional stages.

The net effect is a shift in where computation is spent: more passes through the small model, fewer through the large one.

Why This Matters

Research impact. The paper reframes verification as a cost that should be scheduled adaptively rather than paid unconditionally after every drafting stage. It also identifies a structural limitation of parallel draft models that prior adaptive draft-length work did not address, which points at a concrete gap for follow-up work.

Real-world applications (these follow from faster, output-preserving LLM inference; the abstract itself reports only speedup and lossless decoding):

  • Interactive chat and assistants, where latency per generated token is what users perceive.
  • Code completion and code generation tools, which run generation repeatedly inside an editor loop.
  • Long-form generation such as document summarization or report drafting, where the total number of generated tokens dominates cost.
  • On-device or locally hosted inference, where the compute budget is fixed and every saved target-model pass matters.
  • Agentic and multi-step pipelines, which chain many generation calls and accumulate per-call latency.

Industry relevance. The method is presented as a layer that improves existing speculative decoding methods (EAGLE-3, DFlash, Domino, DSpark, MTP modules) rather than replacing them, and the reported gains are in wall-clock speedup with unchanged outputs — the two properties serving systems care about most. A public code release is planned at https://github.com/naver-ai/DLoop.

Future Directions

  • Choosing the stopping criterion. The abstract describes looping "while the draft model remains confident" but not how confidence is measured or thresholded; designing and validating that signal is an open design question.
  • Generalizing loop-aware training. Whether exposing the draft model to its own unverified hidden states transfers across different draft architectures and target-model pairings is not settled by the abstract.
  • Combining with adaptive draft length. DLoop and adaptive draft-length methods both decide how much drafting precedes verification; how the two interact, or whether they can be unified, is left open.
  • Characterizing the compute tradeoff. Because DLoop deliberately spends more draft-model passes, the conditions under which this remains a net win — model size ratios, batch sizes, hardware, serving setups — would need fuller treatment than the abstract provides.
  • Integration into production serving stacks. Turning a decoding-loop change into deployable throughput and latency gains raises questions about batching, memory, and scheduling that the abstract does not address.

Target Audience

Researchers and engineers working on LLM inference efficiency, particularly those already familiar with speculative decoding, draft-model training, and parallel drafting frameworks such as EAGLE-style methods. It is also relevant to systems and serving engineers who care about latency in deployed generation, and to graduate students studying decoding-time acceleration techniques. Readers without background in speculative decoding will find the abstract's framing of draft/verify passes the minimum prerequisite for following the argument.

Authors’ abstract

Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at https://github.com/naver-ai/DLoop.

Read the original paper