Skip to content
AI.info

Research

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

Overview Research area: LLM agent security, specifically indirect prompt injection and the tokenizer-level mechanics of chat-template forgery. Technical level: Advanced. The paper assumes familiarity

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
arXiv
2609.35932
Published
2026-09-28
Authors
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao

AI summary

Overview

Research area: LLM agent security, specifically indirect prompt injection and the tokenizer-level mechanics of chat-template forgery.

Technical level: Advanced. The paper assumes familiarity with chat templates, byte-level BPE tokenization, reserved/special tokens, input embedding vectors, and serving stacks such as vLLM.

Scope: A measurement study showing that a forged chat-template marker such as <|im_start|> derives most of its attack authority from its reserved token id and learned input vector rather than from its text, using byte-identical prompt pairs on InjecAgent, AgentDojo, and a census of Hugging Face tokenizer configurations.

What This Paper Is About

An LLM agent reads untrusted tool output into the same token stream that carries its own system prompt and role markers, so a malicious tool result can imitate those markers. Prior work shows that wrapping an injected payload in the model's chat template is highly effective, but it changes the marker's visible text and its token ids together, so the two cannot be told apart. This paper holds the bytes fixed and varies only whether the forged markers reach the model as reserved control tokens or as ordinary subwords, to measure how much of the attack's authority comes from the reserved token's learned representation.

Key Contributions

  1. A fixed-byte measurement design. The authors construct prompt conditions that decode to identical bytes and differ only in whether forged markers keep their reserved ids (Reserved) or are encoded as ordinary subwords (Split), plus a Matched control that keeps the reserved ids while adding exactly the same number of extra tokens. The paired difference Δ = ASR(Matched) − ASR(Split) is defined as the identity gap.

  2. Quantification across four open-weight families. The identity gap is 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B, and 8.1 pp on Qwen3-8B direct harm, and carries over to AgentDojo's multi-turn execution-level scoring.

  3. Localisation of the effect to a single learned vector. Vector-replacement experiments show that the mean of the marker's subword vectors does not reproduce the reserved vector, that another reserved control token's vector keeps the attack near full strength, and that on Llama-3.1 the vector of the nearest ordinary token restores the attack in full.

  4. An audit of the existing mitigation's coverage. The Hugging Face option that re-encodes special tokens as ordinary subwords applies only to tokens a configuration declares special; in 33 of 67 distinct tokenizer configurations covering 255 of the 400 most-downloaded chat models on Hugging Face, it leaves tool-protocol tokens intact.

Main Findings

  • Reserved representations carry most of the attack. On 400 paired cases per configuration, Matched stays within 1.7 pp of Reserved on every configuration, so the extra tokens alone cost the attacker almost nothing. The Split condition, with the same bytes and token count, is far weaker: Llama-3.1-8B direct harm falls from 98.2% (Reserved) to 39.7% (Split), below its own plaintext rate of 51.6%. GLM-4.5 direct harm falls from 71.5% to 11.1% (gap +60.0 pp), and data stealing from 92.3% to 27.2% (gap +65.5 pp). Seed-OSS-36B gives gaps of +58.6 pp (direct harm) and +39.3 pp (data stealing).

  • Qwen3-8B is the exception, for a specific reason. Its identity gap is +8.1 pp on direct harm and +0.5 pp on data stealing, because the text of the markers carries most of the attack there: the split template beats plaintext by 55 pp on direct harm and 48 pp on data stealing. Qwen3-8B emits a reasoning block before acting; suppressing it widens the direct-harm gap from 8.1 to 49.8 pp, with Split falling from 75.3% to 35.8%.

  • A published null result reflects a ceiling, not an absent effect. Deng et al. (2026) split role markers into single characters on Qwen3-8B and saw the target tool call move from 100.00% to 99.99%. On the authors' data the same split costs the attacker 54 pp of successful episodes on that model, and their readout exceeds 0.99 on every Qwen3-8B case of their conditioned subsample, leaving no room to fall. The authors report that reading the data that way reproduces the near-zero shift on Qwen3-8B while Llama-3.1 drops by 40 points.

  • The effect is not about where the extra tokens land. Move the extra tokens to the ordinary text ending at the marker or starting after it, and the gap shifts by at most 1.4 pp; splitting after the marker lowers the gap by more than 3 pp only on Llama-3.1, where it stays above 40 pp. Using a per-character split rule instead gives a larger gap of 54.2 pp on Qwen3-8B direct harm, and all 21 gaps across three combinations of split rule and matched control are positive.

  • The authority sits in one vector at the marker position. On 510 direct-harm cases, replacing only the input vector at each reserved marker position with the nearest ordinary token's vector restores the attack on Llama-3.1 (98.4% versus 98.2% with the reserved vector), while the mean of the marker's subword vectors reaches only 58.4%. Replacing it with another reserved control token's vector (<|endoftext|> on Qwen3-8B, <|python_tag|> on Llama-3.1) keeps the attack at or near full strength on both models. On Qwen3-8B neither ordinary vector recovers the gap.

  • Embedding proximity alone is not sufficient. Restoring the reserved id is worth 18 to 47 pp on three models, whereas moving the split closer to the reserved vector in embedding space is worth at most 17 pp and nothing on Llama-3.1.

  • Instruction tuning increases the reserved marker's authority. Across three base and instruction-tuned pairs (Qwen3-1.7B, Qwen3-8B, Seed-OSS-36B), measured on the logit margin of the attacker's tool over the user's tool, the instruction-tuned checkpoint has the larger gap in every pair, while the cost of the extra tokens does not shift.

  • The gap carries into multi-turn agent tasks. On AgentDojo's held-out split of 409 task pairs spanning all four suites, scored by the benchmark's execution-level security check, the gap is 10.3 pp on Qwen3-8B and 6.8 pp on Qwen3-32B. On a second three-suite split that adds Llama-3.1 and Seed-OSS-36B, the gap is positive in all ten combinations of model and suite.

  • The standard tokenizer mitigation misses the tool channel. In 33 of 67 distinct tokenizer configurations among the 400 most-downloaded chat models on Hugging Face, tool-protocol tokens such as <tool_call> and <tool_response> are declared outside the option's reach. A forged block built only from those tokens yields identity gaps of 9.4 to 19.9 pp on every model the authors test that declares them.

  • An adaptive attacker recovers most of the attack. Against six respelling rules fixed in advance, the defence removes up to 51.8 pp of the attacker's best success rate. But a searched spelling (115 to 133 candidates per model, selected on separate calibration cases) comes within 1.6 to 12.2 pp of the undefended attack, and on three of the four models the best spelling replaces each marker with an ordinary token near the reserved one in embedding space.

Methodology in Plain English

The attacker controls the bytes of a single tool result and nothing else; the defender controls the tokenizer call, and with it which token ids those bytes become. Both sides see the same string in every log.

For each InjecAgent case, the authors build prompts that share the content, injection site, user task and tool set, and differ only in how the forged markers inside the payload are encoded. Reserved uses the attacker's payload under default tokenization. Split encodes each marker with the ordinary vocabulary, which is what the standard mitigation produces. Matched keeps the reserved ids but forces the same number of extra tokens, case by case, out of ordinary text elsewhere in the tool response, so Split and Matched share bytes and token count and differ only in whether the markers keep their reserved ids. Plaintext is the same injection without template markers.

Splitting a marker does two things at once: it removes the reserved ids and it adds tokens (between 7 and 33). Matched separates those two effects. The authors note that Matched is a conservative control, because it splits ordinary words into pieces the tokenizer would never produce, a perturbation models are known to tolerate (Zheng et al. report Qwen-2.5-7B retaining 93.4% of its performance under random re-segmentation), so any cost of that unusual segmentation falls on Matched and would shrink the measured gap.

Each configuration runs three to five times at temperature 0 on the same cases, and a gap counts as established only if an exact paired test finds it significant in every run. Before any generation, every case is checked that Reserved and Split decode to identical bytes, that Split contains no reserved id while Reserved and Matched do, and that Matched has exactly Split's token count. Models are served with vLLM from raw token ids and decoded greedily. Token budgets are set per family to keep truncation rare, not tuned on attack success: 1536 tokens for Qwen3, Llama-3.1 and GLM-4.5, and 4096 for Seed-OSS-36B. GLM-4.5 is served in FP8.

The paper evaluates Qwen3-8B, Llama-3.1-8B-Instruct, GLM-4.5 and Seed-OSS-36B-Instruct, with Qwen3-32B as a scale check, sampling 400 cases for each of InjecAgent's two attack types, direct harm and data stealing, giving eight configurations. A later section adds AgentDojo, which runs complete multi-turn tasks and counts an attack only when its security check finds the injected goal reached after the tools have executed. Vector-swap experiments intervene directly on the model's input embeddings rather than on bytes an attacker can send.

Why This Matters

The paper reframes chat-template prompt injection as a tokenizer-level decision rather than purely a syntax-imitation problem. If most of an injection's authority comes from the reserved token's learned input vector, then defences that operate on text — Unicode look-alikes, string filters, renaming markers — address the surface rather than the mechanism, and evaluations that report only text cannot see the difference.

Real-world applications:

  • Self-hosted agent deployments. Any stack where the deployer controls tokenization (the paper uses vLLM 0.11.0, confirming that control-token strings placed inside user content or a tool return reach the model as reserved ids for six chat and tool markers on both channels) can use the fixed-byte contrast to estimate how much of an attack the tokenizer option blocks before relying on it.

  • Model and tokenizer distribution. Because the standard mitigation only covers tokens a configuration declares special, publishers of tokenizer configurations need to declare, or otherwise handle, the tool-protocol tokens through which agent frameworks pass untrusted tool output.

  • Agent framework design. Frameworks that pass tool results through <tool_call> and <tool_response> markers leave an open channel in 33 of 67 distinct tokenizer configurations covering 255 checkpoints.

  • Safety evaluation of injection defences. The contrast applies unchanged to models trained to resist injection, such as StruQ, SecAlign and Meta SecAlign, where it can test whether a defence removes this authority or only the surface cues that trigger it.

Industry relevance. The findings touch the serving layer (tokenization is a server-side choice), the model distribution layer (Hugging Face tokenizer configurations), and the agent layer (tool-output channels). The authors argue that defences should be judged by the token ids they let reach the model, and that studies of template attacks should report token ids alongside text.

Future Directions

  • Test whether injection-resistant training removes the authority or only the surface cues. The authors explicitly propose applying the fixed-byte contrast to StruQ, SecAlign and Meta SecAlign.

  • Extend beyond self-hosted open-weight models. Hosted APIs that accept only strings and reject reserved-token strings in user content (for example through tiktoken's disallowed_special) are stated as outside the threat model, so whether the effect generalises to observable hosted deployments is open.

  • Harden or replace the tokenizer-side mitigation. The source-aware encoder described in the appendix, which tokenizes trusted and untrusted spans separately with a tokenizer whose special-token matching step is removed, requires no training and round-trips byte-exactly, but needs validation beyond the configurations served here, including the constrained variant needed when reserved ids lie inside the base vocabulary, as Kimi-K2's five tool-protocol ids do.

  • Address the adaptive attacker. Because removing the reserved id removes the default attack's advantage but not the authority itself, an open question is how to defend against an attacker who searches for ordinary tokens that sit near the reserved vector in embedding space, a channel that recovered 92.2% on Llama-3.1, 70.4% on GLM-4.5 and 72.6% on Seed-OSS-36B in the defended condition.

Target Audience

This paper is most useful to LLM security researchers working on prompt injection and agent robustness; engineers who own self-hosted inference and tokenizer configuration for agentic systems; maintainers of agent frameworks that pass untrusted tool output through protocol markers; and safety evaluators who need a measurement that separates

Authors’ abstract

Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as &lt;|im_start|&gt; can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged markers as subwords, with the text held fixed and a control for the extra tokens this adds, lowers attack success on the InjecAgent benchmark by 39 to 66 percentage points on three of four open-weight families, and the gap carries over to multi-turn agent tasks in AgentDojo. On Qwen3-8B the gap is 8 points, because without reserved ids the model still recognises the forged turn from its text by reasoning; suppressing the reasoning block widens the gap to 50. The authority sits in the single learned vector at the marker position: the mean of the marker's subword vectors does not reproduce it, the vector of the nearest ordinary token restores the attack on Llama-3.1, and an adaptive attacker who searches for non-reserved markers finds such embedding neighbours on three of four families. In every base and instruction-tuned pair we test, instruction tuning strengthens the model's preference for reserved markers. The standard mitigation, a tokenizer option that encodes special tokens as ordinary subwords, applies only to tokens a configuration declares special, so in 33 of 67 distinct tokenizer configurations, covering 255 of the 400 most-downloaded chat models on Hugging Face, it leaves intact the tool-protocol tokens through which agents read untrusted tool output, and the gap persists on that channel.

Read the original paper