Research
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
Overview Research area: Large language model (LLM) agents, failure analysis, and agent post-training (supervised fine-tuning, preference learning). Technical level: Advanced. The paper assumes familia

- arXiv
- 2609.40111
- Published
- 2026-09-30
- Authors
- Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li, Heng Ji
AI summary
Overview
Research area: Large language model (LLM) agents, failure analysis, and agent post-training (supervised fine-tuning, preference learning).
Technical level: Advanced. The paper assumes familiarity with agent harnesses, rollout traces, supervised fine-tuning (SFT), preference optimization (DPO), verifier-based evaluation, and train/holdout protocol design.
Scope: The paper introduces the Agent Error Dataset (AED), a collection of 50,228 error–diagnosis pairs drawn from text-based agent systems, together with a five-stage construction pipeline (AET) and experiments on correction testing, diagnosis training, and actor repair training.
What This Paper Is About
When an LLM agent fails a task, the failed rollout contains more information than a single reward signal: what the agent observed, which actions it chose, and how the environment responded. The core problem is that a low reward does not say which decision should be revised or what action should replace it. The authors build AED to connect natural failures with trace-cited diagnoses, proposed corrections, and—where replay is possible—recorded execution evidence, then use those records to train diagnostic and acting policies.
Key Contributions
-
A scaled collection of natural agent failures with provenance. AED contains 50,228 error–diagnosis pairs linked to 38,278 stored source-trace blobs and 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems (1.31 pairs per blob on average). Source traces and execution metadata are retained so failures can be re-diagnosed without repeating the original rollout.
-
The Agentic Error-to-Training (AET) pipeline. A five-stage pipeline that (1) collects natural failures, (2) diagnoses the error location, responsible agent, evidence-based explanation and proposed correction, (3) grounds the diagnosis in the student-visible trace through structural and semantic checks, (4) replays corrections where supported against a fresh original-action retry from the same checkpoint under matched policy, harness, budget and verifier settings, and (5) builds objective-specific training views for diagnosis SFT, recovery SFT and action preferences.
-
Controlled measurement of correction utility. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points (task-clustered 95% interval: 28.4–37.0) over original-action retries.
-
An empirical study of error-aware post-training. Full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout, where the strongest prompted reference in that comparison scores 54.7%.
Main Findings
- Corrections beat retries in matched replay: First-proposal corrections raise verifier success from 18.4% to 51.1% (a 32.7-point gain) versus original-action retries from the same checkpoint under matched settings, without selecting among repeated proposals.
- Success is not the same as benefit: Of the 3,062 paired attempts, the first proposal succeeds on 1,564, but the original-action retry also succeeds on 470 of them (30.1%). The correction-only and retry-only cells contain 1,094 and 93 attempts respectively, and 1,405 attempts fail in both arms—so the net paired gain comes from discordant cells, not from counting every successful correction as improvement.
- Diagnosis production yield: The citation-first judge produces trace-cited diagnoses for 93.5% of the common failure pool, compared with 48.4–65.8% for the evaluated alternatives. The authors state this measures production yield, not independent label accuracy.
- Supplied-location assistance: In a separate study, diagnosis-guided continuation scores 31.3%, versus 26.0% for generic reconsideration and 18.6% for replay. The paired gain over generic reconsideration remains uncertain. Both studies test correction at a supplied location and neither tests autonomous detection or actor post-training.
- Internal diagnosis learning: Full-diagnosis SFT raises exact-step agreement at each of four nested training-set sizes; even the smallest subset improves on the untrained base. The strongest prompted reference in this comparison scores 54.7%, against 63.6% for the trained student averaged over three seeds on the 943-case holdout.
- Public benchmark transfer is environment dependent: Under a unified protocol, the 1,656-task arm's three seeds improve on base on Who&When, but after excluding flagged task overlap the interval includes zero; mean TrajErrBench accuracy remains below base. Under a different public protocol (948-task, seed-17 student), full-diagnosis training loses both responsible-agent and exact-step accuracy on Who&When, and an answer-format continuation recovers part of the deficit while changing training exposure as well as format.
- Frontier reference comparison: Under the same frozen cases and first-call contract, the answer-format student exceeds four prompted frontier references on responsible-agent attribution in both hand-crafted conditions, but not on the algorithm-generated conditions, exact step, or AgentErrorBench.
- Repair training is environment dependent: In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training. WebShop-lite gains accompany shorter episodes, while many newly lost TextQuest tasks hit the step limit. Action-only and reflective WebShop gains and all TextQuest losses survive a post-hoc Holm-adjusted multiplicity correction across tasks; other differences are not detected. The evaluated repair-containing recipes score higher on WebShop-lite and lower on TextQuest.
- Real-environment recovery is mixed: On real ALFWorld and ScienceWorld, the preventive arm loses to the untrained policy. Only that recipe has this real-environment comparison.
- Historical failure profiles differ by harness: In an earlier trajectory-deduplicated taxonomy study, 2,092 of 2,604 trajectories were classified; the nine displayed groups contain 2,588 trajectories (2,086 classified), with 16 falling outside these groups. Action-error labels lead in 8 of the 9 shown harnesses, including the planner–executor harness (61.7%), while the pi agent instead concentrates labels on verification and observation (61.3%).
- Diagnoses revisit the same location: Across 405 multi-round attempts, 1,064 of 1,744 adjacent diagnoses retain the same attributed step (61.0%); 363 move later and 317 earlier.
- Collection composition is broad but uneven: BFCL supplies 11.0% of pairs, the five largest environments supply 44.0%, and the five largest harness families supply 78.8%. Multiple diagnoses add 11,950 pairs beyond one per source blob (9,846 source blobs have multiple diagnoses). Pairs per source task have a median of 2 and a 90th percentile of 10.
- Human agreement: The table lists AED at 85.5% raw-step human agreement (reported as 59/69), with the caveat that metrics differ across resources.
Methodology in Plain English
The authors begin from failures rather than successes. A run counts as failed when its environment adapter reports an unsuccessful terminal outcome within the allowed budget, and infrastructure or grader faults are screened out. Each retained record links a failed action–observation trace to one free-text diagnosis (attributed step, responsible agent, trace-cited explanation, and proposed correction), plus review decisions and any replay branches.
Construction runs through five stages: collect failures across environments, harnesses and policies; diagnose the error location and responsible agent; ground the diagnosis in the trace the student can actually see through structural and semantic checks (recording both accepted and rejected proposals); replay the correction where supported against a fresh original-action retry from the same checkpoint under matched policy, harness, budget and verifier settings; and finally build separate training views. Diagnosis records need no replay, recovery targets require an executed passing continuation, and preference pairs require comparable executed alternatives from the same state with a positive outcome contrast.
The views themselves use standard post-training objectives. Diagnosis examples pair the failed trace with the diagnosis target, and loss is applied only to target assistant tokens (context and tool observations receive no loss), normalized per example. Preventive examples pair pre-error history with an executed passing action; post-error variants add the original erroneous action and its rejection as masked context before supervising the correction, with or without a diagnosis-derived reflection. Actor targets use normalization spanning all ranks and accumulation steps. The authors evaluate SFT, and state that an offline action-DPO pilot establishes no recovery benefit.
For evaluation, trainable models start from Qwen3-8B, and each study uses a frozen, objective-specific population. Internal splits hold out source tasks; missing or unparseable responses count as misses. The collection itself, the replay cohort, the separately frozen diagnosis release (4,319 rows over 2,309 source tasks in 15 environments, each environment produced by up to nine harness families and ten policy models), and the objective-specific training subsets all have different admission rules and their counts are not additive. Human review involved four paper authors with doctoral or AI/LLM research backgrounds, reported with anonymous rater identifiers and aggregate statistics.
Why This Matters
Impact on research. Most agent post-training uses outcome rewards that do not specify which decision to revise. AED supplies paired failures, diagnoses, proposed corrections and replay evidence, letting researchers study diagnosis and action learning on a common collection of failures rather than across resources with incompatible execution settings, annotation targets and correction evidence. It also contributes a methodological point: successful corrections alone do not isolate corrective benefit, so successful retries and failed corrections are both retained.
Real-world applications.
- Software and repository agents that inspect code, run commands and must recover from mistaken edits or commands.
- Web and shopping agents where a wrong click or search must be repaired within a step budget.
- Tool-using assistants that must recognize when a tool call was rejected and pick a valid alternative.
- Failure triage systems that group recurring error modes across different agent harnesses to guide which errors to prioritize.
Industry relevance. Teams deploying agents need to know whether a failure is the model's fault or the harness's, and whether a proposed fix actually works under the same execution conditions. AED's matched-replay protocol and same-checkpoint original-action controls give a template for measuring correction utility in production-like settings, and the reported environment dependence of both diagnosis transfer and repair training warns against assuming that a fix validated in one environment will help in another. The collection is text-only and covers benchmarks, synthetic tasks and simplified ports, so environment counts are not counts of public benchmarks.
Future Directions
- Autonomous detection and recovery. Both correction studies test a supplied error location; neither tests whether an agent can detect its own failure or an actor policy trained to recover autonomously.
- Broad, overlap-controlled transfer. Public benchmark gains shrink or disappear after excluding tasks shared with training, and the authors note that internal holdouts share environments and often harnesses and policies with training. Cleaner transfer evaluation across disjoint environments remains open.
- Recipe isolation for repair training. The single-seed actor comparisons differ in task pools, exposure and optimizer updates, so the contrasts measure the combined recipe rather than repair supervision alone; reflection adds no detected benefit over action-only targets, and budget effects are not separated from exposure or update-count differences.
- Broader modality and coverage. AED is text-only, coverage has sparse coding cells and procedural stand-ins, and the authors report that environment counts are not public-benchmark counts. Successfully re-diagnosing stored failures without new rollouts also opens the question of whether added diagnoses improve training.
Target Audience
Researchers and engineers working on LLM agents, agent failure analysis, and agent post-training (SFT and preference optimization) will get the most from this paper. It is also relevant to teams building evaluation harnesses and verifiers for interactive agents, to dataset builders who need a template for pairing failures with diagnoses and controlled replay evidence, and to practitioners who need to decide whether a proposed fix for an agent failure is actually better than simply retrying the original action.
Authors’ abstract
An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems. We retain source traces and execution metadata to support cross-setting failure analysis and re-diagnosis without repeating the original rollout. Our five-stage Agentic Error-to-Training (AET) pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence. Where replay is supported, we compare corrections with original-action retries from the same checkpoint under matched execution settings. We then construct separate training views for diagnosis and actor recovery. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points. Using a separately frozen diagnosis release, full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout. The strongest prompted reference in this comparison scores 54.7%, and mean agreement improves at each of four increasing training-set sizes. In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.