Research
How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression
Overview Research area: Mechanistic interpretability of agentic large language models, specifically the internal mechanism behind the call-or-no-call decision (whether a model invokes an external tool

- arXiv
- 2610.09624
- Published
- 2026-10-07
- Authors
- Xijie Gong, Tingxu Han, Jiahao Zhang, Wei Song, Ziqi Ding, Hanqi Yan, Youcheng Sun, Lijie Hu
AI summary
Overview
Research area: Mechanistic interpretability of agentic large language models, specifically the internal mechanism behind the call-or-no-call decision (whether a model invokes an external tool or answers directly).
Technical level: Advanced. The paper assumes familiarity with activation patching, residual-stream analysis, direct logit attribution (DLA), and Transcoder-based feature decomposition.
Scope: The paper isolates a single residual-stream direction in Qwen3-8B that causally controls the first-token tool-call decision, traces how that direction is formed by suppressive features activated by "analysis" request verbs, and tests whether the same mechanism recurs across seven models.
What This Paper Is About
Agentic LLMs receive long, heavily scaffolded prompts that mix role instructions, tool schemas, format templates, and a user request, and they must decide whether to emit a tool call or a direct text response. Because these prompts contain hundreds of tokens with many entangled components, prior mechanistic work on short prompts of 10–30 tokens offers no obvious single variable to manipulate. The paper's goal is to find one controllable variable in agentic prompts, use it to build matched prompt pairs, and then locate and explain the internal state that carries the call-or-no-call decision.
Key Contributions
-
A controlled causal interface for agentic prompts. The authors show that swapping a single request verb, an execution verb such as write for an analysis verb such as discuss, reliably flips the first generated token between
<tool_call>and direct text, while the scaffold, task body, and available tool stay fixed. They build 500 such paired prompts across Python, Java, and C++ (300 for mechanistic analysis, 200 held out for evaluation), drawing on MBPP, APPS, HumanEval, and CodeContests. -
A causally sufficient and necessary tool-call vector. A vector, written μ_Δ (mu sub Delta), estimated as the mean execution-minus-analysis activation difference at the layer-24 prediction position in Qwen3-8B, is both sufficient and necessary for the decision. The paper reports that it generalizes beyond the discovery prompts to new domains, to native multi-turn τ²-Bench trajectories, and to verb-free requests.
-
A formation account centered on suppression. Rather than execution verbs actively generating a call signal, the evidence indicates the prompt scaffold installs a tool-call prior and analysis verbs suppress it through Transcoder features that signal tool use is unnecessary, written against μ_Δ across layers 21–23.
-
A readout account and cross-model replication. Scaffold-reading attention heads and late MLP features read the resulting state into the
<tool_call>opening, and the same mechanism is reported to recur across seven models from the Qwen, Mistral, and Granite families.
Main Findings
-
One verb flips the decision. Matched prompts differing only in the request verb produce a reliable behavioral contrast: execution prompts yield
y = <tool_call>, analysis prompts do not. The paper uses five execution verbs (add, build, ...) and five analysis verbs (discuss, explore, ...). -
The decisive state localizes to layer 24 in Qwen3-8B. Activation patching at the prediction position recovers
<tool_call>on 100% of held-out analysis prompts at L24, with the patching effect shifting from the verb position to the prediction position across layers. -
μ_Δ is causally sufficient and necessary. Adding μ_Δ to analysis prompts raises the top-1 call rate from 0.00% to 100.00% and the mean tool-call logit from 25.11 to 32.50 (up 7.39), giving a normalized sufficiency score of 1.03. Removing μ_Δ from execution prompts lowers the top-1 call rate from 100.00% to 0.00% and the logit from 32.28 to 24.80 (down 7.48), giving a normalized necessity score of 1.04.
-
Transfer across tools, domains, and action wording. With the coding-derived direction held fixed at L24 and rescaled to the target-domain norm, vector addition induces calls on 100.0% of corrupt prompts and removal suppresses calls on 100.0% of clean prompts across web retrieval (
web_search), SQL execution (run_sql), and email dispatch (send_email), using 100 held-out pairs per domain. The mean normalized logit shift is 0.88 for addition and 0.94 for removal. A norm-matched random vector changes no top-1 decisions. -
Transfer to native multi-turn trajectories. On τ²-Bench Telecom trajectories of 5,385–15,690 context tokens with 16–43 available tools and 200 decision points per arm, at 1× gain vector removal suppresses 36.7% of baseline tool calls and addition induces calls on 70.0% of turns with a baseline text response. A norm-matched random vector switches 1.5% and 0.0% of these turns. Removal decreases the mean
<tool_call>logit by 5.55 and addition increases it by 19.42. -
Transfer to verb-free requests. Across 600 non-imperative requests, among 160 evaluated requests with a baseline Qwen3-8B tool call, removal suppresses 93.1% of calls at 1× gain and 100.0% at 1.5× gain, compared with 16.2% for a norm-matched random vector at the same gain (mean paired logit change −9.95).
-
The scaffold installs a request-sensitive call prior. On 200 held-out tasks, neutral requests produce a mean first-token call probability of 0.8486 versus 0.0038 for analysis requests. The format template alone drives call probabilities to 1.0000 (neutral) and 0.9952 (analysis), removing it nearly eliminates calling (6.32 × 10⁻⁷ and 8.31 × 10⁻¹⁰ with role instructions plus tool schema), and removing the tool schema weakens the separation (0.9521 versus 0.4108). The task body alone yields a 32.3% call rate; execution requests reach 100% and analysis requests fall to 0%.
-
Formation is MLP-dominated and suppressor-driven. The layer-20–23 "formation window" accounts for a sharp rise in the projected gap; MLPs supply 77.7% of the projected write versus 22.3% for attention heads, with MLP23 contributing the largest single write at 22.5 on the g_l scale. Analysis-active features dominate L21–L23 (K_corrupt/K_clean of 1.34, 1.39, and 4.18 respectively), with L23 labeled "analysis-verbs."
-
A large analysis-non-execution feature family. The largest family by contribution contains 115 labelled analysis-non-execution features with a mean held-out |κ| of 0.0633 versus 0.0156 for the 292-feature execution-request family; its total is 7.28 within a full-feature total of 14.34, and the window-level K_corrupt/K_clean ratio is 1.99.
-
Suppressor ablation changes behavior. Zeroing the prediction-position writes of the five training-selected highest-|κ| suppressors shifts the held-out analysis-side L24 input state by +5.30 along μ̂_Δ, closes 8.82% of g_l, raises the mean call margin by +1.38, and raises the call rate from 3.0% to 25.5%, with 22.5% strict recovery (45/200). A layer-matched random suppressor control shifts the state by +0.14 (0.23% gap closure) with 0.5% strict recovery.
-
Mediation test recovers clean calls. Replacing execution-side L20–L23 prediction-position MLP outputs with paired analysis-side outputs, then restoring only the lost μ̂_Δ projection at L24, recovers the baseline gap and 100% of clean calls (200/200).
-
Readout through scaffold-reading heads. L29H9, L33H11, and L33H29 shift toward the tool-format template and away from role instructions on execution prompts, with L33H29 showing the largest attention-head DLA increase of +13.45. Adding μ_Δ raises L33H29's DLA from 11.13 to 24.43.
-
A late structural MLP feature. MLP34 has the strongest single-component patching effect in the readout sweep, yielding a 41.0% call rate on 200 held-out feature-evaluation pairs. Transcoder feature F109925 responds to tool-schema boundaries; adding μ_Δ raises its mean activation from 67.6 to 126.0 (projected call-token write from 3.35 to 6.25), close to execution-side values of 124.4 and 6.17. Replacing only its analysis-side activation yields a 37.0% call rate versus 6.5% for the non-structural control F91365.
-
Replication across seven models. Full-state patching recovers calls on 93.5–100.0% of analysis prompts; vector interventions give normalized sufficiency scores of 0.81–1.03 and necessity scores of 0.61–1.04. K_corrupt > K_clean in every model, preserving the suppressive asymmetry. Localization layers vary: Qwen3-4B L26, Qwen3-8B L24, Qwen3-14B L34, Qwen3.5-4B L31, Qwen3.5-9B L31, Mistral-Small-3.2-24B L25, Granite-3.3-8B L35.
-
Structural parallel to refusal. The authors describe the mechanism as structurally mirroring refusal: a strong prior is present, and a compact suppressor overrides it.
Methodology in Plain English
The researchers needed a way to change exactly one thing in a long agentic prompt. They found it in the request verb: asking the model to write a function tends to trigger a tool call, while asking it to discuss the same function tends to produce text. Everything else in the prompt, the role instructions, tool schemas, format templates, and task body, stays identical within a pair, so any behavioral difference can be attributed to that one word.
Because the first generated token is a clean measurable output (either the tool-call marker or ordinary text), the call-or-no-call decision becomes a single-token prediction problem. The authors use activation patching to find where in the network that decision lives: they copy a residual activation from an execution prompt into an analysis prompt and check whether the model now emits <tool_call>. Sweeping layers and positions localizes the decisive state to layer 24 at the final prompt position in Qwen3-8B.
They then average the difference between execution and analysis activations at that state over the 300 training pairs to get the tool-call vector, μ_Δ. Adding it to analysis prompts and subtracting it from execution prompts tests sufficiency and necessity. To explain where the vector comes from, they project component outputs onto its unit direction, decompose MLP writes with Transcoders into interpretable features, and score each feature by how differently it activates across the two prompt types and by how strongly its decoder direction aligns with the vector. Downstream, they use direct logit attribution and single-component patching to identify which attention heads and MLP features turn the resulting state into the first token. Finally, they repeat the pipeline on six additional models using model-specific prompt templates and call-opening tokens.
Why This Matters
Impact on research. The paper pushes mechanistic interpretability from short, synthetic prompts into the long, scaffolded prompts that real agents actually receive, and it extends causal analysis to a behavior (action initiation) rather than a factual recall or a refusal. It connects behavioral tool-use benchmarking to an internal causal account, and it reports that the same functional organization, a scaffold-induced prior overridden by suppressive features, holds across seven models even though the specific layers and components differ.
Real-world applications:
- Coding assistants that must decide between writing a file through a tool and explaining code in text, the exact setting the discovery prompts use.
- Web-search agents where unnecessary calls add latency and cost, and missed calls leave questions unanswered.
- Database agents that execute SQL only when the request actually requires it rather than on every analysis-style question.
- Email-dispatch agents where writing a message must be distinguished from discussing what a message should say.
Industry relevance. Because the vector transfers without re-estimation to long multi-turn trajectories and verb-free requests, it suggests a cheap steering direction for controlling when deployed agents act. Separating tool selection and argument correctness from the call-or-no-call decision also gives a cleaner debugging target for agent reliability. The paper explicitly notes it does not evaluate broader tool-use correctness such as whether a call selects an appropriate tool, produces valid arguments, or completes the task.
Future Directions
- Extending the formation analysis beyond the behaviorally screened coding pairs and fixed scaffold, since the paper states its applicability to other prompt constructions remains to be established.
- Resolving the fine-grained computation that maps scaffold information and request semantics onto the vector, which the paper says it does not yet resolve.
- Determining whether explicit and implicit (verb-free) requests share identical formation mechanisms, since transfer supports reuse of a downstream call state but not that.
- Evaluating the broader correctness of tool use, including appropriate tool selection, valid argument generation, and task completion, which is outside the paper's scope.
- Explaining why localization layers and component identities vary across models while the suppressive asymmetry (K_corrupt > K_clean) stays constant.
Target Audience
Mechanistic interpretability researchers studying transformer internals; agent and tool-use researchers who want a causal account of action initiation rather than accuracy improvements; model-safety researchers interested in steering agent behavior through residual directions, given the structural parallel to refusal; and engineering teams building production agents who need to control when a model calls a tool versus answers directly. Readers without a background in activation patching, logit attribution, or Transcoder decomposition will find the methods sections demanding.
Authors’ abstract
Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and the user's request across hundreds of tokens, creating a noisy, highly entangled context in which no single controllable variable for mechanistic analysis is obvious. To obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., \textit{write}) with an analysis-verb (e.g., \textit{discuss}) reliably flips the decision, suggesting it is mediated by a compact internal state. We construct 500 such paired prompts across Python, Java, and C++ (300 for mechanistic analysis, 200 held out for evaluation). We trace the decision to a vector, $μ_Δ$, that is both causally necessary and sufficient and generalizes beyond the discovery prompts to native multi-turn $τ^2$-Bench trajectories and verb-free requests. Behavioral ablations show that the scaffold establishes a tool-call prior; Transcoder decomposition then reveals that analysis verbs suppress this prior through features signaling that tool use is unnecessary, whereas execution verbs largely leave it intact. Downstream scaffold-reading attention heads and MLP features read out the resulting state, and the same mechanism recurs across seven models from the Qwen, Mistral, and Granite families. Our code is available at https://github.com/XijieGo/MI4ToolCalling.