UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
Organizations: Independent Researcher
Abstract
Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles. On 12 held-out TEST workflows across two open-weight models, two frameworks, and three recovery paradigms (5,760 executions / 2,880 paired trials) in the frozen lost-acknowledgment study, nominal competence reached 83.54% while conditional recovery success rate (CRSR) fell to 46.72%, with naive retry producing duplicate external effects in 53.33% of trials. Extensions to commercial API models reproduced this competence-recovery separation. Evaluations across complementary execution boundaries show that recovery is phase-dependent: before mutation, methods perform similarly without duplicate effects among capable trials; during partial mutation, naive retry, per-call idempotency, and zero-privilege journaling collapse on the evaluated composite workflows; after commit but before acknowledgment, verification and server-side idempotency substantially improve safety. These findings demonstrate that evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities in autonomous agents.
Figures & tables
| Method | Control | CRSR | EOR | DER | URR |
|---|---|---|---|---|---|
| ( ) | ( ) | ( ) | ( ) | ( ) | |
| : Naive Retry | 802 / 960 | 42.77% | 42.77% | 53.33% | 53.33% |
| : Idempotency | 802 / 960 | 53.49% | 53.49% | 45.42% | 45.42% |
| : EvoUndo-RB1 | 802 / 960 | 43.89% | 43.89% | 51.67% | 51.67% |
| Contrast | — | +10.72 pp | +10.72 pp | -7.91 pp | -7.91 pp |
| Contrast | — | +1.12 pp | +1.12 pp | -1.66 pp | -1.66 pp |
| Task ID | Domain / Operation | Control | CRSR | CRSR |
|---|---|---|---|---|
| RB-PAY-004 | Subscription Renewal | 80 / 80 | 40.0% | 100.0% |
| RB-PAY-005 | Wire Transfer | 80 / 80 | 65.0% | 100.0% |
| RB-DB-004 | Schema Migration | 48 / 80 | 52.1% | 72.9% |
| RB-CRM-004 | SLA Renewal | 80 / 80 | 0.0% | 0.0% |
| RB-STOR-005 | Chunk Swap | 0 / 80 | N/A ∗ | N/A ∗ |
| Model | Control | CRSR | DER | CRSR | CRSR |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | 80.69% | 21.08% | 65.62% | 33.33% | 22.11% |
| GLM-5.2 MaaS | 77.71% | 21.89% | 51.25% | 33.07% | 21.66% |
| Pooled | 79.20% | 21.48% | 58.44% | 33.20% | 21.89% |
| Method | PRE CRSR | DURING CRSR | POST CRSR | DURING DER |
|---|---|---|---|---|
| : Naive Retry | 77.17% | 0.00% | 43.13% | 82.50% |
| : Idempotency | 77.42% | 0.00% | 45.98% | 81.56% |
| : EvoUndo-RB1 | 76.68% | 0.00% | 45.57% | 82.50% |
| : Verify-Retry | 76.92% | 25.89% | 75.00% | 54.37% |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | Specification / Scale |
|---|---|
| Base Workflows | 36 Workflows across 8 Enterprise Domains |
| Split Allocation | DEV (14), VALIDATION (10), TEST (12) |
| Held-Out TEST Tasks | 12 Tasks (Table 6 ) |
| Evaluated Models | Llama-3.1-8B-Instruct, Llama-3.2-3B-Instruct |
| Orchestration Frameworks | Direct Tool Calling (F1), LangGraph Agent (F2) |
| Recovery Paradigms | Naive Retry (B0), Idempotency (B2), EvoUndo (B5) |
| Task ID | Domain | Baseline Workflow Description | Reversibility | Idempotency / Deduplication Support |
|---|---|---|---|---|
| RB-CLOUD-004 | Cloud | Blue-Green Service Deployment Cutover | Compensatable | Routing state overwrite (Inherent idempotent) |
| RB-CRM-004 | CRM | Enterprise SLA Tier Renewal | Compensatable | None (Unkeyed mutable REST API) |
| RB-DB-004 | Database | Schema Migration with Batch Audit Log | Compensatable | SQLite schema unique constraint |
| RB-DB-005 | Database | Read-Modify-Write Bank Dividend Allocation | Compensatable | None (Unkeyed SQL ledger insert) |
| RB-GIT-004 | Git | Release Tag Cut and Cherry-Pick | Reversible | Git tag collision rejection (Inherent unique) |
| RB-MSG-004 | Messaging | High-Urgency Pager Callout with Delivery ACK | Irreversible | None (Unkeyed webhook push) |
| Task ID | Domain | Workflow Name | Pairs | Control Pass | CRSR | CRSR |
|---|---|---|---|---|---|---|
| RB-CLOUD-004 | Cloud | Blue-Green Deployment Cutover | 240 | 100.0% (240 / 240) | 100.0% (80 / 80) | 100.0% (80 / 80) |
| RB-CRM-004 | CRM | Enterprise SLA Tier Renewal | 240 | 100.0% (240 / 240) | 0.0% (0 / 80) | 0.0% (0 / 80) |
| RB-DB-004 | Database | Schema Migration with Audit Log | 240 | 60.0% (144 / 240) | 52.1% (25 / 48) | 72.9% (35 / 48) |
| RB-DB-005 | Database | Bank Dividend Allocation | 240 | 100.0% (240 / 240) | 0.0% (0 / 80) | 0.0% (0 / 80) |
| RB-GIT-004 | Git | Release Tag Cut & Cherry-Pick | 240 | 100.0% (240 / 240) | 100.0% (80 / 80) | 100.0% (80 / 80) |
| RB-MSG-004 | Messaging | Pager Callout with Delivery ACK | 240 | 100.0% (240 / 240) | 0.0% (0 / 80) | 0.0% (0 / 80) |
| Task ID | Domain / Operation | Control ( ) | Observable? | Probe State | CRSR | Recovery Action / Mechanism |
|---|---|---|---|---|---|---|
| RB-CLOUD-004 | Cloud Traffic Switch | 80 / 80 | Yes | PRESENT | 100.0% | Active probe confirms route; suppresses duplicate retry. |
| RB-CRM-004 | CRM SLA Tier Renewal | 80 / 80 | No | UNKNOWN | 100.0% | Zero-privilege abstention; avoids duplicate SLA update. |
| RB-DB-004 | DB Schema Migration | 48 / 80 | Yes | PRESENT | 100.0% | Schema catalog query; suppresses duplicate DDL. |
| RB-DB-005 | DB Blind Audit Append | 80 / 80 | No | UNKNOWN | 100.0% | Zero-privilege abstention; avoids duplicate log insert. |
| RB-GIT-004 | Git Release Tagging | 80 / 80 | Yes | PRESENT | 100.0% | Public Git CLI inspection; suppresses duplicate tag. |
| RB-MSG-004 | Slack Channel Alert | 80 / 80 | No | UNKNOWN | 100.0% | Zero-privilege abstention; avoids duplicate notification. |
| Reversibility Regime | Tasks ( ) | Control Pass ( ) | Naive Retry ( ) | Idempotency ( ) |
|---|---|---|---|---|
| Reversible | 1 task | 100.0% (240 / 240) | 100.0% (80 / 80) | 100.0% (80 / 80) |
| Compensatable | 9 tasks | 78.80% (1,702 / 2,160) | 33.27% (189 / 568) | 48.42% (275 / 568) |
| Irreversible | 2 tasks | 96.25% (462 / 480) | 48.05% (74 / 154) | 48.05% (74 / 154) |
| Execution Regime | Latency (ms) | Prompt Tokens | Completion Tokens | Tool Calls | LLM Invocations |
|---|---|---|---|---|---|
| Control Runs ( ) | 2,924 1,482 | 738.9 112.4 | 87.2 24.1 | 1.54 0.62 | 2.04 0.48 |
| Faulted Runs ( ) | 3,782 1,694 | 746.2 115.3 | 96.4 26.2 | 2.40 0.71 | 2.32 0.65 |
| vs. Control | +29.3% | +1.0% | +10.5% | +55.8% | +13.7% |
| Property | DEV Evaluation Subset (6 of 14 tasks) | VAL Split (10 Tasks) | TEST Split (12 Tasks) |
|---|---|---|---|
| Workflows | 6 Evaluated Tasks | 10 Tasks | 12 Tasks |
| Executions (Pairs) | 720 (360) | 2,400 (1,200) | 5,760 (2,880) |
| Layer A Control Pass | 63.89% (230 / 360) | 74.50% (894 / 1,200) | 83.54% (2,406 / 2,880) |
| Unconditional RSR | 20.28% (73 / 360) | 44.83% (538 / 1,200) | 39.03% (1,124 / 2,880) |
| Conditional CRSR | 31.74% (73 / 230) | 60.18% (538 / 894) | 46.72% (1,124 / 2,406) |
| Framework (F1 - F2) | 0.00 pp | 0.00 pp | +0.50 pp |
| Benchmark / System | Multi-Turn | Post-Mutation | Paired Control/ | Competence | Ground-Truth | Wire-Level | Side-Effect Safety | Pluggable |
|---|---|---|---|---|---|---|---|---|
| Workflows | Ambiguity (Lost ACK) | Fault Design | Conditioning (CRSR) | State Oracles | Effect Logging | Metrics (EOR/DER) | Substrates | |
| ToolBench ( Qin et al., 2024 ) | ||||||||
| SWE-bench ( Jimenez et al., 2024 ) | ✓ | |||||||
| WebArena / AgentBench ( Liu et al., 2024 ; Zhou et al., 2024 ) | ✓ | |||||||
| OccuBench ( Hu et al., 2026 ) | ✓ | |||||||
| ReliabilityBench ( Gupta, 2026 ) | ✓ |
| Model | Recovery Method | Control ( ) | RSR ( ) | CRSR ( ) | 95% CI (CRSR) | DER | MER | |
|---|---|---|---|---|---|---|---|---|
| Gemini 3.8 Flash | : Naive Retry | 480 | 81.04% (389) | 17.29% (83 / 480) | 21.08% (82 / 389) | [2.73%, 43.87%] | 65.62% | 18.33% |
| Gemini 3.8 Flash | : Idempotency | 480 | 80.00% (384) | 27.92% (134 / 480) | 33.33% (128 / 384) | [8.33%, 61.33%] | 53.96% | 20.00% |
| Gemini 3.8 Flash | : EvoUndo-RB1 | 480 | 81.04% (389) | 18.12% (87 / 480) | 22.11% (86 / 389) | [2.22%, 45.65%] | 63.12% | 20.21% |
| GLM-5.2 MaaS | : Naive Retry | 480 | 77.08% (370) | 17.29% (83 / 480) | 21.89% (81 / 370) | [0.00%, 50.00%] | 51.25% | 31.46% |
| GLM-5.2 MaaS | : Idempotency | 480 | 78.12% (375) | 26.46% (127 / 480) | 33.07% (124 / 375) | [3.28%, 65.07%] | 41.67% | 31.87% |
| GLM-5.2 MaaS | : EvoUndo-RB1 | 480 | 77.92% (374) | 16.88% (81 / 480) | 21.66% (81 / 374) | [0.00%, 50.25%] | 51.88% | 31.25% |
| Task ID | Domain | Method | Control ( ) | CRSR ( ) | DER | Cap. Dup. (%) | MER | URR | |
|---|---|---|---|---|---|---|---|---|---|
| RB-DB-004 | Database | 80 | 70 | 0.00% (0 / 70) | 97.50% (78) | 98.57% | 12.50% (10) | 97.50% | |
| RB-DB-004 | Database | 80 | 75 | 0.00% (0 / 75) | 97.50% (78) | 97.33% | 10.00% (8) | 97.50% | |
| RB-DB-004 | Database | 80 | 68 | 0.00% (0 / 68) | 95.00% (76) | 97.06% | 16.25% (13) | 95.00% | |
| RB-DB-004 | Database | 80 | 69 | 0.00% (0 / 69) | 91.25% (73) | 91.30% | 15.00% (12) | 91.25% | |
| RB-STOR-004 | Storage | 80 | 80 | 0.00% (0 / 80) | 100.0% (80) | 100.0% | 0.00% (0) | 100.0% | |
| RB-STOR-004 | Storage | 80 | 80 | 0.00% (0 / 80) | 100.0% (80) | 100.0% | 0.00% (0) | 100.0% |