Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
Organizations: Apodex
Abstract
An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems. We retain source traces and execution metadata to support cross-setting failure analysis and re-diagnosis without repeating the original rollout. Our five-stage Agentic Error-to-Training (AET) pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence. Where replay is supported, we compare corrections with original-action retries from the same checkpoint under matched execution settings. We then construct separate training views for diagnosis and actor recovery. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points. Using a separately frozen diagnosis release, full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout. The strongest prompted reference in this comparison scores 54.7%, and mean agreement improves at each of four increasing training-set sizes. In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.
Figures & tables
| Work | Errors | Envs | Harn. | Models | Records | Step | Evid. | Exec. fix | Ctrl. | Train | Agree. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Who&When ( Zhang et al., 2025b ) | nat. | 2 | 2 | – | 184 | – | – | ||||
| Who&When Pro ‡ ( Liu et al., 2026 ) | inj. | 26 | 15 | – | 12,326 | – | .73 | ||||
| AgenTracer ( Zhang et al., 2026a ) | nat.+inj. | 6 | 6 | – | 2,000 | D | – | ||||
| AEGIS ( Kong et al., 2026 ) | inj. | 6 | 6 | 1 | 9,533 | – | D | .81 | |||
| MAST ( Cemri et al., 2026b ) | nat. | 8 | 7 | 6 | 1,642 | – | .88 | ||||
| TRAIL ( Deshpande et al., 2025 ) | nat. | 2 | 2 | 2 | 841 | – | – |
| (a) Who&When: agent / exact step (%) | ||||
|---|---|---|---|---|
| Model or training resource | HC | HC + gold | AG | AG + gold |
| Qwen3-8B base | ||||
| AED diagnosis SFT | ||||
| answer-format diversity | ||||
| (b) AgentErrorBench: exact step / step+module (%) | ||||
| Model or training resource | ALFWorld | WebShop | GAIA | Env. macro |
| Who&When | |||||
|---|---|---|---|---|---|
| Model or training resource | HC | HC + gold | AG | AG + gold | AEB |
| Prompted references, same unified attribution prompt | |||||
| Gemini 3.7 Flash | |||||
| GPT-5.6 Sol | |||||
| Claude Opus 5 | |||||
| Claude Sonnet 5 | |||||
Appendix figures & tables64 assets
Supplementary material from the paper’s appendix.
Appendix
| Measure | Count |
|---|---|
| Error–diagnosis pairs | |
| Distinct source blobs | |
| Environment-scoped source tasks | |
| Environments / harness families / policy models | |
| Nonempty environment–harness cells | |
| Source blobs with multiple diagnoses |
| Rank | Environment | Pairs | Share (%) | Source tasks |
|---|---|---|---|---|
| 1 | BFCL | 5,521 | 10.99 | 470 |
| 2 | ALFWorld | 4,838 | 9.63 | 1,431 |
| 3 | AgentBench-DB | 4,407 | 8.77 | 1,120 |
| 4 | Mind2Web (static) | 4,050 | 8.06 | 527 |
| 5 | AgentBench-OS | 3,309 | 6.59 | 485 |
| 6 | -bench retail | 3,134 | 6.24 | 324 |
| Rank | Harness family | Pairs | Share (%) | Source tasks |
|---|---|---|---|---|
| 1 | ReAct (single) | 16,646 | 33.14 | 4,622 |
| 2 | Native tool calls | 9,614 | 19.14 | 3,231 |
| 3 | smolagents | 4,614 | 9.19 | 1,899 |
| 4 | AutoGen | 4,368 | 8.70 | 1,610 |
| 5 | OpenHands | 4,329 | 8.62 | 1,651 |
| 6 | LangGraph | 3,873 | 7.71 | 1,406 |
| Population | Counting unit and size | Use in this paper |
|---|---|---|
| Current collection | pairs over source blobs and source tasks | Source-linked collection; separate from the frozen training release |
| Collection inventory | configuration records with steps and a trace reference | Collection coverage; not an admitted training split |
| Panel annotation audit | attempts; pass citation grounding | Annotation validity under automated checks |
| Frozen diagnosis release | rows over source tasks in environments (train, development, holdout); more rows over previously exposed tasks ship as an exploratory split | The evaluated dataset; a separate, earlier snapshot |
| Pilot diagnosis split | train, development, holdout rows | Source for the three-seed pilot ( shared source tasks per arm) |
| Final paired cohort | tasks in families of the -row training split ( tasks); held-out rows over tasks | Final-scale comparison (Table 25 ) |
| Question (majority of three seats) | Majority verdict | Split | Unanimous | Fleiss | ||
| Attributed step is a defensible decisive step | yes | alternative | no | |||
| Quoted evidence supports the claim | supports | partially | does not | |||
| Proposed replacement is plausibly better | yes | unclear | no | |||
| All three seats: step defensible and evidence at least partial of ; step yes and evidence supports of (Wilson to ) | ||||||
| Stage | Train | Development | Holdout |
|---|---|---|---|
| Input records | ( ) | ( ) | ( ) |
| After deterministic checks | ( ) | ( ) | ( ) |
| After semantic filtering | ( ) | ( ) | ( ) |
| Prompt | Both labelled | Mode | Family | Malformed (judge B) |
|---|---|---|---|---|
| Seed | 275 | 0.258 | 0.186 | 88 |
| Rubric | 266 | 0.327 | 0.285 | 104 |
| Our holdout ( ) | AgenTracer test ( ) | ||||
|---|---|---|---|---|---|
| Training resource | Tasks | Seeds | Step | Step | Unit step |
| Qwen3-8B base | — | ||||
| AED (ours) | |||||
| AgenTracer released data v1.0.0 | |||||
| ( AED AgenTracer data), paired | [0.21, 4.81] | [-12.41, -5.70] | [-14.05, -6.96] | ||
| Environment | Rows | Source tasks | Harness families | Policy models | Paired replay attempts |
|---|---|---|---|---|---|
| ALFWorld | |||||
| AppWorld | |||||
| AgentBench-DB | |||||
| BFCL | |||||
| Mind2Web (static) | — | ||||
| -bench retail |
| Environment | Pairs | Orig. pass | Prop. pass | Disc. | McNemar | |
|---|---|---|---|---|---|---|
| BFCL | ||||||
| AgentBench-DB | ||||||
| -bench retail | ||||||
| AgentBench-OS | ||||||
| -bench airline | ||||||
| AppWorld |
| Continuation | Pairs | Orig. pass | Prop. pass | [95% CI] | Disc. | McNemar |
|---|---|---|---|---|---|---|
| Substituted ReAct (single) continuation | [39.5, 48.2] | |||||
| Same harness, state restored | [42.3, 51.7] | |||||
| Multi-agent state partly restored | [44.2, 57.8] |
| Evaluation subset | Cases | Gain (pp) | interval |
|---|---|---|---|
| All cases | |||
| No flagged task overlap | |||
| Flagged task overlap |
| Interface | Cohort | Seed 17 | Seed 202 | Seed 828 | Mean [95% interval] |
|---|---|---|---|---|---|
| Shared | HC | [ ] | |||
| HC + gold | [ ] | ||||
| AG | [ ] | ||||
| AG + gold | [ ] | ||||
| AEB | [ ] | ||||
| TEB | [ ] |
| Model | Clipped policy | Full policy | First call | Continued |
|---|---|---|---|---|
| Base | ||||
| SFT, seed 17 | ||||
| SFT, seed 202 | ||||
| SFT, seed 828 |
| Recipe | Tasks | Rows | Examples | Batch | Epochs | Updates |
|---|---|---|---|---|---|---|
| Diagnosis: seeds 17, 202, 828; sequence limit 24,576 | ||||||
| Full diagnosis | ||||||
| Full diagnosis | ||||||
| Actor: seed 17; sequence limit 12,288 | ||||||
| Success-only | ||||||
| Preventive | ||||||
| Diagnosis configuration | Pairs | Yield | USD / input |
|---|---|---|---|
| AgentDebugX: all-at-once | |||
| AgentDebugX: binary search | |||
| AgentDebugX: counterfactual | |||
| AgentDebugX: heuristic | |||
| AgentDebugX: ensemble | |||
| AgentDebugX: success reference |
| Meter scope | Cells | Logged USD |
|---|---|---|
| Completed cells: rollout | ||
| Completed cells: diagnosis | ||
| Completed cells: combined | ||
| Stopped cells (settled meters) | ||
| Unsettled meters (logged to date) |
| Environment | Runs | Diag. | USD | USD/D |
|---|---|---|---|---|
| AgentBench DB | ||||
| AgentBench OS | ||||
| ALFWorld | ||||
| BabyAI | ||||
| BFCL | ||||
| GridWorld |
| Configuration | Calls/input | USD/input | USD/pair | USD/50k pairs |
|---|---|---|---|---|
| AgentDebugX: all-at-once | ||||
| AgentDebugX: deep analysis | ||||
| Multi-model consensus | ||||
| AED citation-first judge |
| Audit | Recorded observation | Consequence |
|---|---|---|
| Corpus terminal shape | An audit of unique traces selects of ALFWorld failures and of ScienceWorld failures. The separate ALFWorld-lite adapter has of . | Unsuccessful runs end with finish , without recorded parser/tool errors or exhaustion. This proxy does not enumerate self-repair opportunities. |
| Held-out failure events | Text-to-SQL: failures; WebShop-lite: . | No recognized non-infrastructure error event or step-limit termination. Missing event coverage can affect this classification. |
| Lexical subset | Of those failures, text-to-SQL and WebShop-lite tasks contain an earlier observation with a verifier-signal token absent from the task, after stopword filtering. | Fixes task membership for initial-state reevaluation; does not show that the policy saw or recognized a complaint. |
| Training contamination | The strict parser refused list_tables() on of episodes; none of the repair rows was built from that refusal. | Use the corrected parser for all evaluated arms. |
| Environment | Strict (%) | Corrected (%) | Outcome changes |
|---|---|---|---|
| Text-to-SQL | pp; gained, lost / | ||
| WebShop-lite | pp; flipped | ||
| ALFWorld-lite | — | pp; flipped |
| Supervision | Train tasks | Macro | Micro | Unit step |
|---|---|---|---|---|
| Qwen3-8B base | 0 | |||
| Execution-assisted diagnosis | ||||
| Teacher-only diagnosis |
| Benchmark | Cases | Base | Execution-assisted | Teacher-only |
|---|---|---|---|---|
| Who&When HC | ||||
| Who&When AG | ||||
| AgentErrorBench | ||||
| TrajErrBench |
| Continuation at the supplied location | Attempts | Seeds | Success | |
|---|---|---|---|---|
| Original action re-applied | — | ref. | ||
| Generic reconsideration | — | [ , ] | ||
| Diagnosis and proposed replacement | — | [ , ] | ||
| Deployed coach text | — | [ , ] |
| Training tasks | Three seed scores (%) | Mean SD | Source |
|---|---|---|---|
| Run reports | |||
| Run reports | |||
| Run reports | |||
| Run reports |
| Target | Tasks | Seed | Epoch | Seven-field | Compact |
|---|---|---|---|---|---|
| Base | — | — | |||
| Full diagnosis | |||||
| Full diagnosis | |||||
| Full diagnosis | |||||
| Full diagnosis | |||||
| Full diagnosis |
| Protocol and split | Base | Full FT | Multiformat | |||
| Training format: AgentErrorBench | ||||||
| Training format: Who&When AG | ||||||
| Training format: Who&When AG, WG | ||||||
| Training format: Who&When HC | ||||||
| Training format: Who&When HC, WG | ||||||
| Training format: TrajErrBench |
| Committed (arm on W1) | Re-decoded, same box as base | |||||
| Protocol and split | Holm | Holm | Contract | |||
| Training format: AgentErrorBench | ||||||
| Training format: Who&When AG | ||||||
| Training format: Who&When AG, WG | ||||||
| Training format: Who&When HC | ||||||
| Training format: Who&When HC, WG | ||||||
| Training rows | Seed | Correct / cohort | Exact (%) | Format fails | vs. base |
| TrajErrBench | |||||
| 0 | — | 91 / 486 | 18.72 | 0 | — |
| 948 | 17 | 86 / 486 | 17.70 | 0 | 0.6570 |
| 1,656 | 17 | 75 / 486 | 15.43 | 0 | 0.0888 |
| 1,656 | 202 | 75 / 486 | 15.43 | 1 | 0.0805 |
| 1,656 | 828 | 102 / 486 | 20.99 | 0 | 0.2664 |
| Checkpoint | Pairs | Prefers chosen (%) | Implicit reward (%) | Mean margin (nats) | Dev prefers chosen (%) |
|---|---|---|---|---|---|
| Qwen3-8B base | — | — | |||
| DPO seed 17 | |||||
| DPO seed 202 | |||||
| DPO seed 828 | |||||
| Mean sd over seeds |
| Comparison | Observed result | Scope |
|---|---|---|
| Paired replay | Discordant pairs: replacement wins, losses. Source-harness stratum: pairs, points . | Selected replayable attempts; the control reruns the original action. No diagnosis-release row joins this replay study. |
| Located recovery | Deployed coach versus clean diagnosis: points . | Supplied error location, untrained debugger; the interval does not establish a framing benefit. |
| Compact-target pilot | Three-seed micro: execution-assisted, teacher-only, base. | Construction arms differ in teacher access and engine. All six paired X–J intervals include zero. |
| Reasoning control | Base thinking-off: exact step, unit+step on cases. | One run. A single gold unit makes unit+step sensitive to output naming. |
| Public positional control | Test-selected constant steps score , , , . | Diagnostic upper envelope for constant-step guesses; not a deployable baseline. |
| Wave 1 | Wave 2 | Wave 3 | Pooled | |
| Diagnosed attempts | 109 | 170 | 228 | 507 |
| structurally ineligible | 0 | 0 | 130 | 130 |
| certificate-eligible | 109 | 170 | 98 | 377 |
| Certificates | 60 | 74 | 15 | 149 |
| first earned in round 1 | 45 | 40 | 9 | 94 |
| first earned in rounds 2–7 | 15 | 34 | 6 | 55 |
| Environment | Type | Source tasks | Pairs | Frozen release |
|---|---|---|---|---|
| ALFWorld | Public benchmark | yes | ||
| AgentBench-DB | Public benchmark | yes | ||
| AgentBench-OS | Public benchmark | yes | ||
| BFCL | Public benchmark | yes | ||
| ScienceWorld | Public benchmark | yes | ||
| -bench retail | Public benchmark | yes |
| Witness class | Environments | Basis |
|---|---|---|
| Observed agreement | BFCL, ScienceWorld, SWE-bench Verified, -bench airline, -bench retail, Terminal-Bench, AgentBench-DB, AgentBench-OS, SWE-smith, Mind2Web (static), ALFWorld, AppWorld | 10 by oracle prefix, 2 (ALFWorld, AppWorld) by recorded trace |
| Measured divergent (excluded) | MCPMark | Two replays of the same recorded prefix disagreed on the verifier verdict |
| No deterministic replay | WideSearch, GAIA | Live web interactions do not support deterministic prefix replay |
| Audit | Measurement | Consequence |
|---|---|---|
| Policy request | Reconstructed manifests cover traces; lack a system turn and have an unmatched digest. | External-framework system turns were re-rendered; reconstruction is distinct from native request logging. New collection records the first request. |
| Rules invoked by diagnoses | On model-reviewed pairs, weighted shares are visible rules, system-only rules, unstated rules, and contradicted rules. | Model estimates, not human validation. Historical debugger inputs omitted policy system prompts and tool lists. |
| Field clipping | Cap: characters. release rows contain a clip; of text is removed. | The debugger saw clipped text while citation matching used the unclipped render. Seven of unresolved citations recover under a wider scorer render. |
| Rediagnosis | On clipped rows, clipped–clipped step agreement is ; clipped–unclipped is . | Aggregate change is comparable to rerunning. The deterministic voter agrees with its unclipped version on ; rows move only without clipping. |
| Recipe | Input response | Evidence and experimental scope |
|---|---|---|
| Diagnosis SFT | Failed trace attribution, optionally with rationale, evidence and correction | Main diagnosis experiments; the label need not identify a unique cause. |
| Preventive actor SFT | Pre-error history actions from a passing continuation | Executed correction and branch lineage; mixed with successes in Section 5.3 . |
| Post-error actor SFT | History, erroneous action and rejection correction, with or without reflection | Eligible state-preserving rejections; other repairs retain preventive targets. Section 5.3 . |
| Action DPO | Shared text prefix chosen and rejected actions | Historical offline pilot only; it does not meet the current strict lineage criterion (Appendix I.15 ). |
| Construction | Target tokens/epoch | Updates (2 epochs) | Updates (3 epochs) |
|---|---|---|---|
| Execution-assisted | |||
| Teacher-only |
| View | Input and target | Admission boundary | v5 development partition |
|---|---|---|---|
| Diagnosis SFT | Failed trace attributed step, explanation, citations, proposal | Resolvable evidence; intentionally unattempted replay allowed; target semantics independent of replay | rows; reference assistance marked in metadata |
| Debugger SFT | Failed trace recorded attribution or recovering intervention | Stricter replay-state gate; reference-assisted labels excluded by default | rows; recovery language only with a passing bundled branch |
| Repair SFT | Pre-action prefix passing continuation’s assistant turns | Executed passing branch; erroneous action and failed suffix excluded | rows in the replay package’s training split; masked prefix and tool observations |
| Outcome | Prefix and candidate replay-derived outcome | Typed execution outcome and direct branch lineage | rows; offline prediction, not a demonstrated RL result |
| Action DPO | Shared prefix with chosen/rejected actions | Valid executed lineage and positive observed preference under the stated protocol | rows; action-specific causal eligibility reported separately |
| Dimension | ADP ( Song et al., 2026 ) | AED |
|---|---|---|
| Unit of data | Successful action/observation trajectory. | Natural failure, diagnosis, proposed fix and available replay branches. |
| Action typing | API, code, message actions; text or web observations. | Same three exports; content map retains speaker, step index, event ID. |
| Provenance | Source dataset name. | Environment/version; harness/version; policy model; run ID; engine version. |
| Failure handling | Success-only corpora; no failure, reward, or outcome field. | Natural failures, never injected; outcome, reward, verifier signal stored. |
| Step-level labels | None. | Blamed step, blamed agent, explanation, alternative steps. |
| Evidence grounding | None. | Verbatim quotes typed by source region/match status; 53.1% of quotes ground the trajectory. |
| Model / condition | W&W-HC | W&W-AG | AEB |
|---|---|---|---|
| ( ) | ( ) | ( ) | |
| Context | |||
| Best constant step (oracle) | |||
| Prompted GPT-5 mini § | |||
| Frozen native-context protocol, no ground truth | |||
| Qwen3-8B base | |||
| Stratum | Records | Preferred step | Single step | Verdict agree. | Joint accept |
|---|---|---|---|---|---|
| Filter survivors | 48 | (88.6%) | (93.9%) | (66.7%) | (62.5%) |
| Judge rejected | 16 | (71.4%) | (77.8%) | (43.8%) | (37.5%) |
| Gate rejected | 16 | (90.9%) | (100.0%) | (68.8%) | (31.3%) |
| All records | 80 | (85.5%) | (92.0%) | (62.5%) | (51.3%) |
| Decision | Response options |
|---|---|
| Diagnosis content | Accept : key content is acceptable; revise : changes are needed; reject : key content is unsupported; insufficient information : the available record does not support a decision. |
| Error localization | Single step ; multiple defensible steps with one preferred step and at least one alternative; no single-step attribution ; no agent error ; or insufficient evidence . |
| Submission | Explicitly confirm the judgment for each record. Step selection is restricted to events in that record. Save the JSON backup and CSV export. |
| Example | Instruction |
|---|---|
| A wrong query at step 0 is corrected at step 2, which returns ; step 3 submits . | Select step 3; exclude the fully recovered query. |
| A budget of 30 is reduced by an irreversible wrong purchase of 20 at step 0; the required item costs 20. | Select step 0; finding the right item later does not restore the budget. |
| A valid read-only query returns HTTP 503 and the harness stops, with no evidence of an alternative tool. | Do not force an agent-error attribution. |
| A required image is unavailable and the grader reports only failure. | Select insufficient information. |
| Step 0 computes and step 1 submits 13 without correction. | Prefer step 0; step 1 may be an alternative. Propagation is not recovery. |
| The agent guesses a PIN; the correct value appears only in later feedback. | Identify the unsupported guess. A correction cannot use the future value; search first only if a search tool was available. |
| Continuation from the checkpoint | Pass (%) | Gain over retry (pp) |
|---|---|---|
| Original-action retry | — | |
| AET: first/only correction |
| (a) Prompt-only baselines | ||
|---|---|---|
| Model | Macro | Micro |
| Qwen3-8B | ||
| GPT-6 Astra | ||
| Claude Sonnet 5 | ||
| Gemini 3.8 Flash | ||
| Claude Opus 5 | ||
| Supervision | Text-to- SQL | ALFWorld lite | TextQuest | WebShop lite | GridWorld | Warehouse |
| Qwen3-8B base | ||||||
| Success-only | ||||||
| + preventive repair | [-1pt] | [-1pt] | [-1pt] | [-1pt] | [-1pt] | [-1pt] |
| + post-error actions | [-1pt] | [-1pt] | [-1pt] | [-1pt] | [-1pt] | [-1pt] |
| + actions and reflection | [-1pt] | [-1pt] | [-1pt] | [-1pt] | [-1pt] | [-1pt] |
| Planned tasks |
| (a) Admission ablation: one candidate pool | |||
|---|---|---|---|
| Admission rule | Retained yield | Tasks / envs | Cost / row |
| No quality checks | (ref.) | / | |
| Grounding checks only | ( ) | / | |
| AET: grounding + semantic review | ( ) | / | |
| Reference | [95% CI] |
|---|---|
| Qwen3-8B base | [ , ] |
| GPT-6 Astra, prompted | [ , ] |
| Claude Sonnet 5, prompted | [ , ] |
| Gemini 3.8 Flash, prompted | [ , ] |
| Claude Opus 5, prompted | [ , ] |
| same target, nested subset | [ , ] |
| (a) Who&When source prompt: agent / exact step (%) | ||||
|---|---|---|---|---|
| Model or training resource | HC | HC + gold | AG | AG + gold |
| Qwen3-8B base | ||||
| AED diagnosis SFT | ||||
| answer-format diversity | ||||
| (b) AgentErrorBench unified prompt: exact step (%) | ||||
| Model or training resource | ALFWorld | WebShop | GAIA | Env. macro |
| Descriptive episode statistics; not a causal mechanism | ||||
|---|---|---|---|---|
| Environment | Error-bearing steps, un-repaired policy | Step limit | mean steps | success (pp) |
| WebShop-lite | ||||
| GridWorld | ||||
| TextQuest | ||||
| Text-to-SQL | ||||
| ALFWorld-lite | ||||
| Environment | Preventive | Action-only | Reflective | R/A | Session gap |
|---|---|---|---|---|---|
| WebShop-lite | 0.0436 | 0.000821 | 0.000535 | 1 | 2.67 |
| GridWorld | 0.21 | 0.0574 | 0.804 | 0.146 | 1.00 |
| TextQuest | 7.43e-06 | 7.03e-05 | 2.43e-06 | 0.627 | 1.67 |
| Text-to-SQL | 0.531 | 0.735 | 0.295 | 0.09 | 3.35 |
| ALFWorld-lite | 1 | 0.607 | 1 | 0.727 | 0.67 |
| Warehouse | 0.607 | 0.424 | 1 | 0.454 | 0.00 |