Research
AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
Overview Research area: Efficient LLM inference — specifically retrieval-based speculative decoding applied to coding-agent pipelines, implemented in vLLM. Technical level: Advanced. The paper assumes

- arXiv
- 2610.01108
- Published
- 2026-10-01
- Authors
- Sumin Lee, Sukmin Cho, Suengjae Lim, Youngjin Kwon
AI summary
Overview
- Research area: Efficient LLM inference — specifically retrieval-based speculative decoding applied to coding-agent pipelines, implemented in vLLM.
- Technical level: Advanced. The paper assumes familiarity with speculative decoding, suffix automata/suffix trees, repository-level coding agents, and serving-level throughput measurement.
- Scope (one sentence): The paper proposes AgSpec, a framework layered on top of existing retrieval engines that supplies coding-agent-specific corpora and adaptive draft-length policies, and measures its throughput gains on SWE-bench Verified, TeamBench, LiveCodeBench, and Terminal-Bench across three models.
What This Paper Is About
Coding agents tackle one user request as a multi-turn session of model calls and tool executions, often split across specialized agents such as a coder that edits files and an executor that runs tests. These agents repeatedly reproduce code, logs, and earlier attempts, which makes retrieval-based speculative decoding — drafting tokens by copying continuations of existing text — a natural fit. The paper's core problem is that existing retrieval-based drafters only capture part of that reusable text (they miss the ongoing session and repository, and store files in a form that differs from what the agent emits), and they size drafts with fixed caps or match-derived rules that ignore how accept length differs across agents and drifts across turns. The goal is a framework that fills these policy gaps without changing the underlying retrieval engine, so any compatible engine can be accelerated as is.
Key Contributions
-
Coding-agent-aware retrieval organized by source. AgSpec maintains three separate corpora — a session corpus holding the current session's prompts, tool outputs, and generated tokens; a workspace corpus holding repository artifacts opened during the session; and a fixed global corpus of static reference text assembled before serving. Live corpora (session and workspace) are built within the session and discarded when it ends.
-
Emission-form indexing of workspace artifacts. Because a unified-diff patch prefixes kept lines with a space and removed lines with a minus, AgSpec indexes each opened artifact in full, plus one copy with every line prefixed by a space and one with every line prefixed by a minus. For scaffolds that write whole files through a tool call, it adds a further copy with newlines and quotes escaped.
-
Source-preference rule for match selection. The engine reports longest match lengths from each corpus (m_S, m_W, m_G). The longer live match wins, with ties going to the session; the global corpus is drafted from only when m_G > max(m_S, m_W) + δ, since its text is unrelated to the repository and a long match there is likely coincidental. Within a corpus, occurrences are grouped by their next three tokens and the most frequent group is chosen (ties broken by most recent occurrence), while the workspace corpus uses the first occurrence.
-
Agent-aware draft-length control. Draft length at step i is K_i = clip(round(α^(a_i) · m_i), 1, K_max^(M,a_i)). The cap K_max is profiled offline from the replayed trajectories used to build the global corpus, as the last position k where Pr(R ≥ k | M, a) ≥ λ, the acceptance probability at which one more draft position covers its verification cost. The scale α starts at 1.0 and is updated online per agent from verification feedback: α^(a_i) ← α^(a_i) + η(τ − I[A_i < D_i]), which the paper frames as a stochastic subgradient step on a quantile loss, targeting full acceptance with probability 1 − τ where τ = 1 − λ. Updates are skipped when a fully accepted draft reaches the cap, since the cap censors the observation.
Main Findings
-
Throughput on repository-level benchmarks: AgSpec achieves the highest or second highest throughput in all evaluated settings, reaching 2.27–4.37× the throughput of autoregressive decoding at batch size 1 and 1.08–4.76× at batch size 16. On average its throughput is 18.0% higher than the fastest prior method. The largest gains occur with Devstral at batch size 1 on SWE-bench Verified, at 1.36× higher than the fastest prior work.
-
Engine-agnostic gains: With SuffixDecoding as the engine, throughput rises by 16.5% on average and accept length by 18.5%; with SAM-Decoding, throughput rises by 64.6% and accept length nearly doubles. While SAM-Decoding trails SuffixDecoding in all twelve settings, AgSpec-SAM outperforms it in nine.
-
Specific model/benchmark result: With SAM-Decoding as the retrieval engine, AgSpec raises the speedup of Gemma3-27B on SWE-bench over autoregressive decoding from 2.72× to 3.61× at batch size 16.
-
Corpus ablation (trace-driven simulation of AgSpec-SAM on SWE-bench Verified): Replacing a per-call corpus with the session corpus raises accept length by 1.6–2.1×, the largest single gain. Adding the workspace corpus contributes 26–47% on its own and still 2.2–7.0% on top of the session corpus, bringing AgSpec to 2.9–4.8× the accept length of the global corpus alone.
-
Emission-form indexing (coder accept length per turn): Indexing artifacts in emission form raises the coder's accept length by 13–39% in the first turn, before its own patches enter the session corpus. From the second turn on, those patches already supply the emitted form, so the gain falls to 3% or less by the fifth turn.
-
Draft-length control (SWE-bench Verified, Devstral, SAM-Decoding engine): At batch size 1, a fixed K = 32 is 14% faster than K = 8, but at batch size 16 it is 30% slower, as rejected tokens consume compute shared by the whole batch. AgSpec is the fastest at both batch sizes, 25% and 16% above K = 8. K = 32 attains a higher accept length than AgSpec at batch size 1 (6.0 vs. 5.5) and a slightly lower accept length at batch size 16 (10.3 vs. 10.4), yet remains slower at both because of the drafts it rejects.
-
Rejection behavior at batch size 1: K = 32 rejects 11–26 tokens per step, whereas AgSpec rejects at most 6, close to K = 8, while still accepting 83–89% as many tokens as K = 32. AgSpec lengthens drafts only as reusable text accumulates, from about 10 tokens in early turns to 21 in the fifth.
-
Isolated effect of the online scale (batch size 16, on top of fixed K = 32 without per-agent caps): The online scale alone raises throughput by 66%, as the share of rejected draft tokens falls from 78% to 50%. In the example coder call, the drafter without it spends about 26 tokens per step over the first 160 steps but has fewer than 2 accepted; with it, drafts shorten to about 4 tokens there and expand to the full draft length once the coder starts copying reusable text, rejecting 81.1% fewer tokens while taking only 11 more steps.
-
Generalization without a repository (LiveCodeBench): Across the three models, the faster AgSpec variant reaches 4.87–7.94× the generation throughput of autoregressive decoding and exceeds the fastest prior speculative method by 35–145%. Its accept length also exceeds that of the prior method with the highest accept length. The workspace corpus is unavailable here, so gains come from the global and session corpora plus the adaptive draft-length controller.
-
Generalization without explicit agent roles (Terminal-Bench): The faster AgSpec variant improves throughput over the fastest prior speculative method by 9.3–27.2%, while its accept length is 16.7–34.3% higher than that of its corresponding retrieval-engine baseline. Here AgSpec uses a shared cap and a single α table instead of role-specific caps.
-
Baseline comparison: AgSpec is compared against autoregressive decoding, five retrieval-based SD methods (PLD, REST, FastCoder, SuffixDecoding, SAM-Decoding), and EAGLE-3 on two repository-level benchmarks.
Methodology in Plain English
The authors keep the retrieval engine itself untouched and instead redesign the policies wrapped around it, so the framework applies to any engine that can search multiple corpora and accept a per-step draft length.
First, they split reusable text by where it comes from. Text produced inside the current session (prompts, tool outputs, generated tokens) goes into a session corpus shared across all calls and turns. Files the session opens go into a workspace corpus, indexed not just as the files stand but in the forms the agent emits them — most importantly the space-prefixed and minus-prefixed line forms that unified-diff patches require. Text unrelated to the session goes into a fixed global corpus built before serving. The retrieval engine queries all three with the same suffix of the decoding context; the longer live match wins, and the global corpus must beat the live match by a margin of δ = 5 tokens before it is used.
Second, they control how many tokens get drafted. They replay the same trajectories that were collected to build the global corpus and record, at each retrieval step, how many retrieved tokens actually matched the recorded continuation. The distribution of that reusable length per agent yields the probability that the k-th draft token would be accepted; the offline cap is the last position where that probability still exceeds the break-even threshold λ, below which verifying another position costs more than it is expected to return. On top of this cap, each agent carries a scale that starts at 1.0 and is nudged after every verification step — increased when a draft was fully accepted, decreased when it was not — so the requested length tracks the current session rather than the agent's history. Some engines, such as SuffixDecoding, have their own draft-length policy; AgSpec uses the smaller of the two values.
Evaluation runs in vLLM over two retrieval engines: SAM-Decoding's suffix automaton (AgSpec-SAM) and SuffixDecoding's suffix tree (AgSpec-Suffix). Both use the same corpus configuration, controller, per-agent caps, and workspace scope, differing only in the underlying retrieval structure. Throughput excludes prefill, and accept length is reported alongside it. Global corpora are built from disjoint tasks — 70 corpus-building and 30 evaluation instances for SWE-bench Verified, 213 and 30 for TeamBench — and held fixed during measurement, with session corpora reset between sessions.
Why This Matters
Impact on research. The paper reframes retrieval-based speculative decoding for agents as a policy problem rather than a matching problem. It shows that the largest accept-length gains come from decisions about what a corpus holds and how long drafts should be, not from a better matcher, and it demonstrates that a framework layered on an unmodified engine can raise that engine's throughput substantially (64.6% average throughput gain with SAM-Decoding). The offline-profiled cap plus online quantile-style scale is a general recipe that other domains with recurring structure and role-dependent output could reuse.
Real-world applications:
- Accelerating AI coding assistants that resolve repository-level issues, where a session spans many turns and repeatedly rewrites files the agent has already read.
- Multi-agent software pipelines that divide work among a planner, executor, verifier, or localizer/coder/executor/reviewer/summarizer, where acceptance differs by role.
- Single-agent tool-using workflows such as ReAct-style loops where a single agent alternates between inspecting files, acting, and interpreting feedback.
- Standalone code generation without a repository, where session and global corpora still supply reusable text.
Industry relevance. The gains matter most where serving is compute-bound: at batch size 16 the verification cost of rejected drafts is shared across the whole batch, which is exactly why long fixed draft lengths stop paying off. Reducing user wait time per turn without changing the model's output distribution is directly relevant to any team serving coding agents at scale. The paper also flags a deployment concern: the workspace and session corpora may contain sensitive prompts, repository content, or tool outputs, so deployments should restrict access to each session's corpus, prevent disclosure to other sessions, and discard it at session end.
Future Directions
-
Extending the source-aware corpus model to other scaffolds and formats. Only unified diffs and whole-file tool-call encoding are described in detail; search-and-replace blocks and line-addressed edit commands are mentioned as formats other scaffolds use, but how AgSpec's indexing should adapt to them is not evaluated.
-
Better handling of the censored observation. The online update is skipped when a fully accepted draft reaches the cap, because the cap censors the measurement. More principled estimators for this censored regime could let the controller learn faster.
-
Generalizing the profiling procedure. The offline cap requires replayed trajectories for each model and agent. Whether caps can be transferred or estimated from far less data — and how they should be set for workflows with no explicit roles — remains open; the Terminal-Bench result uses a shared cap and a single α table as a workaround.
-
Applying the framework to non-coding agents. The three mechanisms are motivated by session, workspace, and emission structure specific to coding. Whether the same policy split helps agents in other text-heavy, multi-turn, tool-using settings is not tested.
Target Audience
This paper is most useful to LLM inference and serving engineers who care about decoding latency in production, to researchers working on speculative decoding and retrieval-based drafting, and to builders of coding agents and multi-agent software pipelines who want faster per-turn response times without retraining a draft model. It is also relevant to systems researchers interested in training-free acceleration that composes with an existing engine. Readers need prior familiarity with speculative decoding, repository-level coding benchmarks, and serving-level metrics such as throughput and accept length to get the most from the analysis.
Authors’ abstract
Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format. It bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37$\times$ at batch size 1 and 4.76$\times$ at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.