Research
Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety
Overview Research area: LLM agent safety, with ties to uncertainty quantification and risk-aware decision making in multi-turn tool-using agents. Technical level: Intermediate. The experimental design

- arXiv
- 2510.16492
- Published
- 2025-10-18
- Authors
- Vamshi Krishna Bonagiri, Ponnurangam Kumaragurum, Khanh Nguyen, Benjamin Plaut
AI summary
Overview
Research area: LLM agent safety, with ties to uncertainty quantification and risk-aware decision making in multi-turn tool-using agents.
Technical level: Intermediate. The experimental design is straightforward (prompt variations on an existing benchmark), but readers benefit from familiarity with agent frameworks such as ReAct, POMDP-style trajectory formalisms, and safety/helpfulness evaluation rubrics.
Scope: A systematic evaluation of whether teaching LLM agents to "quit" (terminate a task) improves safety without meaningfully reducing helpfulness, tested across 12 LLMs, three prompting strategies, and 144 high-stakes scenarios in the ToolEmu framework.
What This Paper Is About
LLM agents increasingly act in real-world environments through tools, but they carry a strong bias toward completing tasks even when instructions are ambiguous or dangerous. The authors propose "quitting" — explicitly allowing and instructing an agent to stop and withdraw when it cannot rule out harmful consequences — as a simple, prompt-only behavioral safety mechanism. The goal is to measure whether this improves safety outcomes and at what cost to task helpfulness.
Key Contributions
- A systematic evaluation of quitting behavior across 12 LLMs and three prompting strategies (Baseline, Simple Quit, Specified Quit) using ToolEmu's 144 high-stakes scenarios, 36 toolkits, and 9 risk types.
- A demonstration that strategic quitting via simple prompting yields a favorable safety-helpfulness trade-off: average safety gains of +0.40 across all models and +0.64 for proprietary models on a 0–3 scale, against an average helpfulness decrease of only –0.03.
- Evidence of a "compulsion to act" in agents: merely offering a quit option (Simple Quit) produced only modest average safety gains of +0.17 with a –0.007 helpfulness impact, far below Specified Quit's +0.40, showing that making the option available is not enough.
- A practical takeaway that adding explicit quit instructions to agent system prompts is an immediately deployable safety intervention requiring no retraining.
Main Findings
- Specified quitting drives the largest safety gains. The Specified Quit prompt achieved the highest safety scores for nearly all models, averaging +0.40 safety across all models and +0.64 for proprietary models.
- Helpfulness costs are minimal. Across all models the Specified Quit condition reduced helpfulness by only –0.03 on average; Simple Quit had virtually no helpfulness impact.
- Dramatic individual model improvements. Claude 4 Sonnet's safety score rose from 1.008 (baseline) to 2.223 (specified quit), a gain of 1.215, with a helpfulness change of only –0.015. GPT-4o rose by +0.972, from 0.894 to 1.866.
- Quit rate correlates with safety gains. High quit rates under Specified Quit — 72.22% for Claude 4 Sonnet and 57.64% for GPT-4o — track the largest safety improvements.
- Proprietary models are more responsive than open-weight models. The Claude, GPT, and Gemini series showed far greater sensitivity; Llama 3.1 8B Instruct, Llama 3.1 70B Instruct, and Llama 3.3 70B Instruct had quit rates under 9% even under Specified Quit (0.00%, 7.64%, and 2.08% respectively), and their safety scores improved little.
- Providing the option alone is insufficient. Simple Quit's average safety gain of +0.17 versus Specified Quit's +0.40 suggests agents default toward task completion unless forcefully told to prioritize safety.
- Notable baseline differences. Safety and helpfulness scores varied widely at baseline; for example, GPT-5 baselined at 1.783 safety and 1.963 helpfulness, while Llama 3.3 70B Instruct baselined at 0.272 safety and 0.627 helpfulness.
- Irreversible-action example. In a Bitcoin withdrawal scenario with two ambiguous addresses, the baseline agent assumed an address and proceeded, the Simple Quit agent recognized the ambiguity but still acted, and the Specified Quit agent quit and requested clarification.
Methodology in Plain English
The authors take an existing agent safety benchmark, ToolEmu, and modify what the agent is allowed to do. ToolEmu tests agents on benign but underspecified user instructions inside an LM-emulated sandbox, using an adversarial emulator to construct risky situations. The benchmark contains 144 test cases spanning 36 toolkits (including BankManager, Venmo, EpicFHIR, AugustSmartLock, GoogleHome, and TrafficControl) and 9 risk categories such as privacy breach, financial loss, and physical harm.
The agent's action space is extended with a dedicated quit action, functionally equivalent to outputting "Final Answer" with an explanation of why it cannot proceed safely. Three prompt conditions are compared: Baseline (ToolEmu's standard ReAct prompt with no quit option), Simple Quit (a quit option is added without safety guidance), and Specified Quit (a quit option plus a "MUST quit" directive triggered by four conditions, such as being unable to rule out negative consequences). To avoid biasing the design, the Specified Quit safety prompt was written by a researcher specializing in uncertainty-aware AI who had no knowledge of ToolEmu or the experimental setup.
Each agent trajectory is scored on a 0–3 scale for safety and helpfulness by LLM evaluators, both powered by Qwen3-32B at temperature 0.0. Agent runs also used temperature 0.0. The paper's evaluation harness is based on ToolEmu's LLM-based emulator and evaluator, following prior work.
Why This Matters
Impact on research: This is described as the first systematic evaluation of quitting as a safety behavior in LLM agents, providing baselines for agent safety awareness across 12 models and showing that a prompting-level intervention can outperform the intuition that conservative behavior must hurt utility. It frames quitting as a behavioral proxy for uncertainty quantification in multi-turn settings, sidestepping numerical confidence estimation.
Real-world applications:
- Financial agents executing irreversible transfers, as illustrated by the Bitcoin withdrawal example and toolkits such as BankManager, Venmo, Binance, and TDAmeritrade.
- Healthcare data handling through tools like EpicFHIR and Teladoc, where privacy breaches carry severe consequences.
- Smart-home and physical security systems such as AugustSmartLock, GoogleHome, IndoorRobot, and TrafficControl, where errors can cause physical harm.
- General enterprise tooling including Terminal, GitHub, Gmail, Slack, and Dropbox, where ambiguous instructions can trigger data loss or security compromises.
Industry relevance: Because the intervention is a prompt modification rather than a training pipeline, it can be deployed immediately in existing agent systems. The finding that open-weight models respond weakly also signals a capability gap relevant to teams choosing models for safety-critical deployments.
Future Directions
- Developing a hierarchy of responses beyond the binary proceed/quit choice, such as asking clarifying questions, requesting permission for risky actions, or terminating the task.
- Fine-tuning models to improve quitting behavior, potentially using automated data generation pipelines.
- Testing whether the results generalize beyond ToolEmu to other agent environments and real-world applications, since the current findings are validated only on that benchmark.
- Improving the instruction-following and risk-awareness of open-weight models, which showed very low quit rates even under explicit directives.
Target Audience
Researchers and engineers working on LLM agent safety, agent deployment, or alignment; practitioners building tool-using agents who need a low-overhead safety mechanism; and readers interested in uncertainty quantification and abstention strategies for multi-turn systems. The paper is accessible to those with basic familiarity with LLM prompting and agent loops.
Authors’ abstract
As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While uncertainty quantification is well-studied for single-turn tasks, multi-turn agentic scenarios with real-world tool access present unique challenges where uncertainties and ambiguities compound, leading to severe or catastrophic risks beyond traditional text generation failures. We propose using "quitting" as a simple yet effective behavioral mechanism for LLM agents to recognize and withdraw from situations where they lack confidence. Leveraging the ToolEmu framework, we conduct a systematic evaluation of quitting behavior across 12 state-of-the-art LLMs. Our results demonstrate a highly favorable safety-helpfulness trade-off: agents prompted to quit with explicit instructions improve safety by an average of +0.39 on a 0-3 scale across all models (+0.64 for proprietary models), while maintaining a negligible average decrease of -0.03 in helpfulness. Our analysis demonstrates that simply adding explicit quit instructions proves to be a highly effective safety mechanism that can immediately be deployed in existing agent systems, and establishes quitting as an effective first-line defense mechanism for autonomous agents in high-stakes applications.