Research
Controlled Decoding Attacks on Black-Box LLMs
Overview Research area: AI security and adversarial machine learning (cs.CR), specifically jailbreak attacks that operate at decoding time against large language models. Technical level: Intermediate.

- arXiv
- 2609.36956
- Published
- 2026-09-29
- Authors
- Jesson Wang, Shawn Li, Wei Yang, Franck Dernoncourt, Ryan A. Rossi, Charith Peris, Yue Zhao
AI summary
Overview
Research area: AI security and adversarial machine learning (cs.CR), specifically jailbreak attacks that operate at decoding time against large language models.
Technical level: Intermediate. The paper assumes familiarity with token-level generation, next-token distributions, sampling, and the general idea of safety alignment, but its core contributions are described largely in terms of interfaces and query budgets rather than deep mathematical machinery.
Scope in one sentence: The paper introduces BlindBias, a framework that jailbreaks aligned LLMs through text-only continuation interfaces by reconstructing next-token distributions from repeated samples and intervening only at positions where the evolving response appears risky.
What This Paper Is About
Existing decoding-time attacks can redirect an aligned model's output, but they require model weights or numerical token probabilities, so they do not apply to interfaces that return only sampled text. The paper asks whether this fine-grained control can be retained when the only available access is repeated stochastic sampling plus continuation from an attacker-supplied assistant prefix. The proposed answer is to estimate the next-action distribution from samples, add probability mass for unobserved actions via a prior, and spend that sampling cost only at selected positions.
Key Contributions
-
BlindBias framework. The authors adapt decoding-time residual control to text-only sampling without access to target weights or log-probabilities, retaining the BiasNet residual formulation from the numerical-probability setting while changing how its inputs are obtained and when it is executed.
-
Three coupled components. Sample-Based Distribution Reconstruction supplies the distributional signal from sampled continuations; Risk-Gated Residual Control uses a prefix-risk model to decide when to reconstruct and how strongly to intervene; Speculative Multi-Token Execution requests multi-token drafts and verifies their prefixes locally to amortize calls to the remote target.
-
Empirical evaluation. The framework is evaluated on four target endpoints (GLM-5, Gemini-3.5-Flash, Qwen3-32B, Kimi-K2.5) and three benchmarks (AdvBench, HarmBench, SORRY-Bench), against four black-box prompt-level baselines (PAIR, GPTFuzz, LogiBreak, FlipAttack).
-
Analyses of reconstruction and cost. The paper analyzes reconstruction quality under different priors, cross-family prior transfer, the trade-off between attack effectiveness and intervention frequency, and an exploratory vocabulary-free string action space. Code is released at https://github.com/JessonWong/controlled-decoding.
Main Findings
-
Highest mean score in most comparisons. BlindBias achieves the highest mean score in 20 of 24 comparisons against four prompt-level baselines. Table 1 reports Harm and Harm Info scores per target, benchmark, and metric, with BlindBias using the soft-gated, global-uniform configuration throughout.
-
Largest gains occur on Gemini-3.5-Flash. BlindBias leads both metrics on all three Gemini-3.5-Flash benchmarks. On SORRY-Bench it improves over the strongest baseline by 1.86 Harm points and 1.19 Info points.
-
The advantage is not universal. FlipAttack leads both metrics on GLM-5 SORRY-Bench, and the benefit of sample-based control varies across targets.
-
Harm and informativeness do not always rank attacks identically. On Kimi-K2.5 AdvBench, LogiBreak achieves a higher Harm score while BlindBias achieves a higher Info score. On Qwen3-32B HarmBench, BlindBias leads Harm while PAIR leads Info.
-
Large distributional changes concentrate at a few positions. Figure 2 plots KL divergence between next-token distributions before and after intervention on GLM-5 and Qwen3-32B; most positions show small changes, with occasional large shifts. This motivates selective rather than uniform control.
-
A prior recovers part of the signal lost to finite sampling. On Qwen3-32B with 100 held-out AdvBench prompts, smoothed empirical counts give PPL 11.85 and Harm/Info 1.50/1.26; adding the uniform prior lowers PPL to 7.07 and raises Harm/Info to 3.03/2.16. A global unigram prior improves predictive fit further (PPL 5.81) yet yields lower attack scores (Harm 2.78, Info 1.94) than the uniform prior. Numerical log probabilities, as a reference outside the sample-only setting, reach PPL 3.17, unseen NLL 7.38, Harm 4.06, and Info 3.08.
-
Predictive fit and steering utility are not interchangeable. The unigram-prior result above and the contextual-prior results both show that better prediction does not automatically mean better attack scores.
-
Contextual priors transfer across model families, with a same-family advantage. Table 2 reports PPL, Harm, and Info as averages over 100 AdvBench prompts. The same-family Qwen3-1.7B proxy performs best on all three metrics (PPL 3.86, Harm 3.90, Info 2.83). Foreign proxies SmolLM2-1.7B (3.93, 3.02, 2.12) and Gemma-3-1B (3.94, 3.38, 2.39) both improve Harm over the shuffled prior control (7.11, 2.69, 2.16), while only Gemma also improves Info.
-
Selective control trades effectiveness for fewer interventions. On 100 AdvBench prompts toward Gemini-3.5-Flash, soft gating improves on hard gating while intervening less often: active positions fall from 9.5% to 5.25% and API calls decrease by 36.8%. Compared with ungated control, the soft pipeline reduces average API calls from 4,000 to 283 (a 92.9% reduction) while Harm Score decreases from 3.96 to 3.58.
-
Active-position frequency alone does not establish endpoint query savings, which also depend on sampling, retries, and response length.
Methodology in Plain English
Threat model. The attacker is given a text-only continuation interface. It accepts a user prompt and an attacker-supplied assistant prefix. The attacker may repeat stochastic continuation requests from the same prompt but cannot see target weights, hidden states, or numerical token probabilities. Returned text is mapped into a local tokenizer's vocabulary to define "actions"; these local actions need not coincide with the provider's internal tokens. A separate exploratory setting uses literal returned strings as actions instead of a tokenizer.
Reconstructing the distribution. Because the API does not expose the next-action distribution, the framework draws K = 50 independent sampled continuations at the same chat history, maps each valid response to one action, and counts how often each action appears. A global-uniform prior with total strength κ = 2 is fused with the counts through a symmetric Dirichlet update, which guarantees that every vocabulary item receives positive probability. Sampling uses temperature 1 and top-p = 1, with one sampled action retained per valid response. Invalid or empty responses are refilled until exactly K valid samples are collected, subject to a retry budget.
Deciding when to intervene. At each ordinary step, the target is queried deterministically to produce a base candidate. After a three-token full-strength warm-up, a local prefix-risk model (Llama-3.1-8B-Instruct) scores the candidate prefix and returns a sigmoid risk score. A soft gate converts this into a residual scale using midpoint τ = 0.1, temperature T = 0.05, and an efficiency cutoff s_min = 0.01, giving an effective execution threshold of approximately 0.3298. When the scale is zero, the base candidate is accepted with no extra sampling. Otherwise the framework reconstructs the distribution and applies a learned residual transformation, BiasNet, to the reconstructed log-probabilities; the next action is chosen by argmax in the experiments. Because scoring depends on the candidate prefix rather than the prompt alone, the same prompt can receive different intervention strengths at different steps.
Training the controller. BiasNet uses a 1,024-dimensional, four-hash count-sketch input projection followed by layer normalization, and is trained for 10 epochs with cross-entropy, AdamW, batch size 32, learning rate 3×10⁻⁴, weight decay 10⁻⁴, mixed precision, and seed 42. The training cache is built from 40 instruction–answer records (indices 100–139 of a pinned LLM-LAT harmful-data revision), with eight records deterministically held out. Positions with zero gate scale are excluded from the loss, and the prefix-risk model is frozen during training. The authors note that training uses reference prefixes while inference uses generated prefixes, and that the shared reconstruction and gate rules do not remove this difference.
Reducing repeated target calls. When the gate suppresses the residual for enough consecutive steps (q_min = 2), the target is asked for a multi-token draft of at most L_max = 80 actions in one request. Every draft prefix is scored locally in a minibatch of size 8, and only the longest run of prefixes satisfying the gate rule is committed. The first violating action and everything after it are discarded, and controlled decoding resumes there. The authors are explicit that this verification enforces the gate rule on accepted prefixes and does not imply distribution preservation or token-for-token equivalence with repeated single-action calls, since a multi-token request may produce different candidates even at zero temperature.
Benchmarks, targets, and metrics. AdvBench contains 520 harmful goals paired with affirmative target prefixes. HarmBench uses its 320-behavior text test split, covering standard and contextual behaviors across multiple harm categories. SORRY-Bench uses its 440 base prompts, balanced over 44 fine-grained safety categories. Gemini-3.5-Flash is accessed through Google's native API, and GLM-5, Qwen3-32B, and Kimi-K2.5 through the OpenRouter API. Responses are scored with Harm Score and Harm Info Score, both assigned by Gemini-3.5-Flash using evaluation prompt templates from prior work. Base requests use temperature 0 and top-p = 1, with a total generation budget of 80 local tokens; response caching is disabled and hidden reasoning tokens are rejected.
Why This Matters
Impact on research. The paper shows that withholding numerical probabilities may not by itself close the decoding-time attack surface when an interface permits repeated sampling and continuation from supplied prefixes. It also connects two previously separate lines of work: decoding-time control, which assumes probability access, and prompt-level black-box jailbreaking, which does not. The finding that distributional changes concentrate at a small subset of positions offers a concrete empirical motivation for selective, prefix-conditioned intervention, and the result that reconstruction accuracy and steering utility diverge challenges the assumption that a better predictor of the target distribution is necessarily a better controller.
Real-world applications:
- Safety evaluation and red-teaming of commercial LLM APIs, where weights and log-probabilities are typically withheld but repeated sampling is allowed.
- API design and abuse-mitigation decisions, since the cost profile of this attack depends on sampling limits, prefix-continuation support, retry policies, and rate limiting.
- Assessment of LLM-based sequential decision-making systems in sensitive domains; the paper motivates its concern partly by reference to high-stakes settings such as precision rehabilitation.
- Agent and tool-use safety, since the same sampling-plus-prefix interface is increasingly exposed by agentic deployments.
Industry relevance. The attacks require only a text-only continuation interface, repeated sampling, and assistant-prefix continuation, all of which are common in deployed serving stacks. The measured cost profile matters operationally: selective gating reduced average API calls from 4,000 to 283 in one reported comparison while lowering Harm Score from 3.96 to 3.58, so the effectiveness-versus-query-cost curve is a practical consideration for both attackers and defenders.
Future Directions
- Endpoint-level query efficiency. The paper states that improving endpoint-level query efficiency remains an important direction, and notes that active-position frequency alone does not establish query savings.
- Control beyond tokenizer-based action spaces. Extending reliable control beyond tokenizer-based action spaces is named as an open direction. The empirical string action space in Section 3.4 is presented only as an exploratory stress test; it cannot emit a target token string that never appears in calibration, and the proxy prior retains only the first proxy token of a possibly multi-token string, so no target token IDs or target tokenizer are used.
- Training–inference prefix mismatch. Training uses reference-answer prefixes while inference uses generated prefixes. The paper states that the shared reconstruction and gate rules do not remove this difference, leaving the effect of this distribution shift open.
- Cross-family prior transfer. Contextual priors from foreign families improved Harm over a shuffled control, but only one foreign proxy also improved Info, and the same-family proxy was best on all three metrics. What determines when a cross-family prior remains useful is left unresolved.
Target Audience
This paper is most useful to LLM security and red-teaming researchers, to engineers responsible for the safety and abuse properties of model-serving APIs, and to readers already familiar with jailbreak literature who want to understand how decoding-time control behaves under a strict text-only interface. It is also relevant to policy and governance audiences weighing whether restricting access to token probabilities constitutes a meaningful defensive measure. Readers without background in language model decoding will need to consult the cited prior work on controlled generation, speculative decoding, and prompt-level jailbreaking to follow the design choices.
Authors’ abstract
Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.