Skip to content
AI.info

Research

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

Overview Research area: Empirical evaluation of tool-using LLM agents, specifically fault recovery in distributed, mutating enterprise workflows (cs.SE). Technical level: Intermediate. Scope: The pape

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
arXiv
2610.05622
Published
2026-10-04
Authors
Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah

AI summary

Overview

Research area: Empirical evaluation of tool-using LLM agents, specifically fault recovery in distributed, mutating enterprise workflows (cs.SE).

Technical level: Intermediate.

Scope: The paper introduces UndoBench, a benchmark of 36 base workflows and 36 fault scenarios across 8 enterprise domains that separates an agent's nominal task competence from its ability to recover safely when a tool call fails mid-execution.

What This Paper Is About

Existing agent benchmarks (SWE-bench, ToolBench, API-Bank, AgentBench, WebArena) mostly measure whether an agent can complete a task under clean, unperturbed conditions, so they cannot tell whether an agent that "succeeds" has corrupted an external system along the way. UndoBench injects realistic distributed-system faults—especially lost acknowledgments, where a remote mutation commits but the confirmation never reaches the client—and inspects wire-level side effects rather than just final answers. The goal is to measure, separately, whether an agent can plan a workflow (task competence) and whether it can diagnose and safely repair execution state after a fault (recovery capability).

Key Contributions

  1. Multi-boundary benchmark formulation. A suite of 36 base workflows and 36 fault scenarios across 8 enterprise domains (Cloud, CRM, Database, Git, Messaging, Payments, Storage, Ticketing), evaluating recovery at distinct execution boundaries relative to external state mutation.

  2. Decoupled paired evaluation and CRSR. A counterfactual paradigm pairing every faulted run with an identical-seed control execution, defining Conditional Recovery Success Rate (CRSR) as recovery success conditioned on paired nominal task success.

  3. Dual state and effect-history oracles. Programmatic oracles that verify physical environment invariants, plus wire-level effect monitors quantifying Duplicate Effect Rate (DER), Missing Effect Rate (MER), and Exactly-Once Semantic Effect Rate (EOR).

  4. Phase-dependent empirical characterization. A frozen confirmatory lost-acknowledgment study of 5,760 executions / 2,880 paired trials, plus prospectively specified post-freeze extensions across commercial API models, complementary execution boundaries, and benchmark-wide idempotency.

Main Findings

  • Competence and recovery are distinct dimensions. Nominal task success reached 83.54% (2,406 / 2,880) while end-to-end recovery success (RSR) was 39.03% (1,124 / 2,880) and conditional recovery success (CRSR) was 46.72% (1,124 / 2,406). Nominal success alone substantially overstates faulted reliability.

  • Naive retry creates duplicate real-world effects. Under B0 (Naive Retry), DER reached 53.33% (512 / 960); the paper states naive retry produced duplicate external effects in 53.33% of trials.

  • Recovery-method comparison on the TEST split (N = 2,880 paired trials). B0: Naive Retry 42.77% CRSR, 53.33% DER. B2: Idempotency 53.49% CRSR, 45.42% DER. B5: EvoUndo-RB1 43.89% CRSR, 51.67% DER. All three had identical control performance of 802 / 960.

  • Idempotency's advantage was not statistically established benchmark-wide. The B2 − B0 contrast was +10.72 pp (bootstrap 95% CI [0.00 pp, +24.23 pp], p_adj = 1.00). Gains concentrated in lost-ACK workflows where endpoints support deduplication. Only three TEST endpoints implement idempotency tracking (RB-PAY-004, RB-PAY-005, RB-MSG-005; 100% CRSR, 0.0% DER); RB-PAY-004 improved by +60.0 pp and RB-PAY-005 by +35.0 pp. The B2 − B0 difference of +20.8 pp on RB-DB-004 came from a datastore uniqueness constraint, not idempotency keys.

  • Zero-privilege journaling did not help. B5 − B0 was +1.12 pp (95% CI [−3.15 pp, +6.31 pp], p_adj = 1.00; DER = 51.67%). Without remote state visibility or registered inverse handlers, the local journal could not confirm whether an in-flight mutation had committed.

  • Duplicate severity depends on operation semantics. Under B0, duplicates occurred in 60.0% of trials on RB-PAY-004 and 35.0% on RB-PAY-005, but 0.0% on RB-DB-004 because SQLite's schema unique constraint blocked the second migration. Across the five workflows lacking idempotency or database constraints, naive retry produced 100.0% (400 / 400) duplicate executions.

  • Model differences were unresolved. Across the 10 TEST tasks supported by both models, M1 (Llama-3.1-8B) reached 44.50% CRSR (445 / 1,000) and M2 (Llama-3.2-3B) 54.17% (547 / 1,010); Δ = −9.67 pp, 95% CI [−27.08 pp, +7.73 pp], p = 1.00.

  • Frameworks were statistically equivalent. F1 (Direct Tool Calling) and F2 (LangGraph) had Control 83.54% vs. 83.54% (Δ = 0.00 pp), RSR 39.24% vs. 38.82% (Δ = +0.42 pp), and CRSR 46.97% vs. 46.47% (Δ = +0.50 pp); TOST against a ±10.0 pp margin gave p_TOST < 0.001.

  • Commercial API models reproduced the gap. Gemini 3.8 Flash and GLM-5.2 MaaS showed nominal control of 80.69% (1,162 / 1,440) and 77.71% (1,119 / 1,440). Under B0, control was 81.04% (389 / 480) and 77.08% (370 / 480), while CRSR fell to 21.08% (82 / 389) and 21.89% (81 / 370)—competence–recovery gaps of 59.96 pp and 55.19 pp. B0 DER was 65.62% (315 / 480) and 51.25% (246 / 480). B2 pooled improvement over B0 was +11.73 pp (33.20% vs. 21.48%; 95% CI [+0.07 pp, +33.25 pp], raw p = 0.0176), but Holm-adjusted p_adj = 0.0528. B5 pooled Δ = +0.41 pp, Holm p_adj = 0.9180.

  • Recovery is phase-dependent. On four qualified composite workflows (N = 320 per cell), PRE_MUTATION CRSR was 77.17% (B0), 77.42% (B2), 76.68% (B5), 76.92% (B6); DURING_MUTATION CRSR was 0.00% for B0, B2, and B5, and 25.89% for B6 (DURING DER 82.50%, 81.56%, 82.50%, and 54.37% respectively); POST_MUTATION_PRE_ACK CRSR was 43.13%, 45.98%, 45.57%, and 75.00%.

  • Pre-mutation redispatch is comparatively safe. Across all 12 TEST workflows with commercial API models (N = 3,840 paired trials / 7,680 runs), CRSR was 78.44% (B0, 593 / 756), 78.86% (B2, 597 / 757), 77.78% (B5, 588 / 756), and 77.34% (B6, 587 / 759), with zero duplicate effects among 3,028 nominally capable trials. B0's boundary contrast was 78.44% vs. 21.48%, Δ = +56.96 pp (95% CI [+27.97 pp, +84.67 pp]), and +57.34 pp on common capable support (D_common = 743).

  • DURING_MUTATION collapse is severe. B0 reached 0.00% CRSR (0 / 310; DER 82.50%), B2 0.00% (0 / 315; DER 81.56%), and B5 0.00% (0 / 308; DER 82.50%). Only B6 recovered, at 25.89% (80 / 309); on RB-CLOUD-004 it achieved 100% CRSR and 0% DER using a public probe, but failed on RB-GIT-004 (unremoved lockfile, exit code 128), RB-STOR-004 (multipart upload state hidden from HeadObject, causing duplicate PUTs), and RB-DB-004 (partial table creation diverging from uncommitted catalog metadata).

  • Endpoint deduplication closes most, but not all, of the gap. With benchmark-wide server-side idempotency (B2-K, N = 960 paired trials / 1,920 runs), CRSR reached 97.24% (741 / 762) versus 33.20% for standard B2, with DER = 0.00%; on common capable support (D_common = 738), 98.64% (728 / 738) versus 33.33%, gap closures of 95.87% and 97.97%. Still, 21 / 762 capable failures remained: 17 before any tool call from empty/invalid hosted responses, and 4 after successful duplicate suppression during multi-turn continuation.

  • Integrity audits passed. Capable-pair prefix equivalence was 100% across PRE_MUTATION (3,028 pairs), B2-K (762 pairs), and DURING_MUTATION (1,242 pairs); all 72 / 72 adversarial oracle test cases passed; 12 provider HTTP retries affected 8 executions with zero tool mutations escaping the proxy wrapper.

Methodology in Plain English

The researchers built a sandboxed benchmark of enterprise-style workflows (cloud provisioning, CRM updates, database migrations, Git operations, messaging, payments, storage, ticketing) and gave agents tools to execute them. Every workflow is run twice under the same random seed: once clean (the control) and once with a single injected fault. Comparing the two runs isolates recovery from planning ability. The primary frozen study fixes the fault at the most hazardous point—after the remote system has committed the change but before the client receives the acknowledgment—and runs 12 held-out TEST workflows across two locally hosted open-weight models, two agent frameworks, and three recovery strategies over 20 seeds, producing 2,880 paired trials and 5,760 executions. The evaluator does not just check the final answer; it watches wire-level calls to count how many times a remote resource was actually mutated, and it checks physical environment invariants. A trial counts as recovered only if the task goal is met with no duplicated and no missing effects. Confidence intervals come from task-clustered bootstrap resampling with 10,000 iterations plus Wilson score intervals, with Holm-Bonferroni correction for multiple comparisons and TOST equivalence testing at a ±10.0 pp margin. After freezing the design, the team ran secondary extensions on commercial API models (Gemini 3.8 Flash, GLM-5.2 MaaS) across seeds 2001–2020, on the PRE_MUTATION and DURING_MUTATION boundaries, and with benchmark-wide server-side idempotency.

Why This Matters

Impact on research. The paper argues that a single end-to-end success number conflates two different capabilities, and formalizes CRSR as recovery success conditioned on nominal competence. It also supplies wire-level effect oracles that detect duplicate mutations which terminal-reward evaluation cannot see. The phase-dependent results argue against treating recovery as one scalar skill.

Real-world applications.

  • Payments and subscription billing, where a retried charge after a lost acknowledgment causes a duplicate billing violation (RB-PAY-004, RB-PAY-005).
  • Database schema migrations, where a repeated migration may be blocked or may partially apply (RB-DB-004).
  • Cloud infrastructure provisioning and traffic cutover, where duplicate deployments or duplicate writes can corrupt state (RB-CLOUD-004).
  • Git and object storage operations, where lockfiles and hidden multipart upload state defeat naive retry (RB-GIT-004, RB-STOR-004).

Industry relevance. Production enterprise systems already rely on idempotency, write-ahead logging, and Sagas, while agent runtimes often treat tool execution as an uninspected request-response call. The finding that B2's gains depend on servers actually supporting deduplication tokens suggests that recovery safety is partly a contract problem between agent harnesses and backend services, not purely a model capability problem.

Future Directions

  • Boundaries left unstudied. POST_ACK_PRE_CHECKPOINT was not faithfully realizable in the current synchronous harness, and DURING_COMPENSATION, cascading faults, and larger workflow suites remain open.

  • Multi-fault and adversarial regimes. UndoBench injects exactly one fault per trajectory; cascading, concurrent, or Byzantine failures are left for future work.

  • Statistical power and clustering. The TEST suite has 12 independent task clusters (K = 12); more seeds reduce within-task variance but do not add independent clusters, limiting the resolution of small effects such as the B2 − B0 contrast.

  • Retry policy sensitivity. B0 uses a fixed 3-retry budget with exponential backoff, and sensitivity sweeps across backoff schedules—along with provider-level retry semantics—are named as future work.

Target Audience

Agent-framework and evaluation researchers, reliability and distributed-systems engineers building tool-using agents for enterprise software, and teams designing fault-tolerance contracts (idempotency keys, verification, compensation) between agent harnesses and backend services. Readers need no formal distributed-systems background, but familiarity with benchmarks and tool-calling agents helps.

Authors’ abstract

Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles. On 12 held-out TEST workflows across two open-weight models, two frameworks, and three recovery paradigms (5,760 executions / 2,880 paired trials) in the frozen lost-acknowledgment study, nominal competence reached 83.54% while conditional recovery success rate (CRSR) fell to 46.72%, with naive retry producing duplicate external effects in 53.33% of trials. Extensions to commercial API models reproduced this competence-recovery separation. Evaluations across complementary execution boundaries show that recovery is phase-dependent: before mutation, methods perform similarly without duplicate effects among capable trials; during partial mutation, naive retry, per-call idempotency, and zero-privilege journaling collapse on the evaluated composite workflows; after commit but before acknowledgment, verification and server-side idempotency substantially improve safety. These findings demonstrate that evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities in autonomous agents.

Read the original paper