Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments
Organizations: Florida International University, Miami, FL, USA
Abstract
Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publish no category for that assumption, and a defect beneath a score is present on every rerun. We treat a tool's advertised surfaces as an executable contract, check the implementation against it, and trace each score's provenance through the task files and evaluator code to the verdicts that derive from state a defective tool should have written. Across 34 audited mutating tools in four benchmarks we confirm seven tool defects and one evaluator property at pinned commits. On injected defects the checker raised no false positive in 25 flags, flagged 2 of 5 negative controls, and missed most: in 29 of 33 scored misses a clause covered the defect but no probe revealed it. The checker's own static half, run alone, flags 14 of 17 confirmed sites, so on these findings the dynamic half confirms and traces rather than discovers. Twelve further AgentDojo tools, with six held-out tools and the seven audited first, complete its 25-tool mutating surface, on which at least 5 tools diverge from their advertised surface as our contracts read it, a rate for AgentDojo alone. No gold trajectory reaches either tau2-bench defect; on 1,120 paths built to isolate the telecom defect, a number fixed by construction, the evaluator rewards a refuel of a suspended line and fails the repaired tool. The clearest case is a clinical benchmark whose tool tells the agent each write executed under a documented no-write design its interface does not disclose; its grader takes that message as evidence, so its action success rate records whether a request carried the expected payload, not whether any record changed.
Figures & tables
| Work(s) | Layer it opens | What its design takes as given |
| Agent-Diff [ 18 ] | State delta across a replica API, as the success criterion | The replica API that reports the delta |
| Gao and Zhou [ 16 ] | The grading script, asking whether an outcome is backed by stored evidence | The tool that produced the evidence |
| Tool-Veritas [ 4 ] | Grader verdict against task outcome | The state the tool wrote |
| ToolFuzz [ 27 ] | Tool documentation, via runtime errors and agent responses | The tool’s state transitions, and any evaluator downstream |
| BenchGuard [ 2 ] , Automated Benchmark Audit [ 3 ] , SafeAudit [ 5 ] | Task artifacts, environment configuration, grader and safety coverage | No published category names the tool implementation |
| Agentic Benchmark Checklist [ 24 ] | Reported benchmark practice | No published category names the tool implementation |
| Class | Checker rule (over pre , post , args , result ) | Status |
| Phantom Effect | success signal true advertised effect delta absent | field-observed |
| Unenforced Precondition | precondition predicate false (no error signal state mutated) | field-observed |
| Ignored Argument | post-state invariant under variation of an argument advertised as effective | field-observed |
| Partial Effect | advertised effect predicate holds fails on the same call | field-observed |
| Invariant Break | environment invariant false after a legal call sequence | mutation-only |
| Reset Leak | snapshot after reset initial snapshot | mutation-only |
| Benchmark | Tools | Mut. | Impl. | Classes (HL) |
| MedAgentBench | 3 | 3 | 1 | 2 (1) |
| tau2-bench | 19 | 19 | 3 | 3 (1) |
| AgentDojo | 7 | 7 | 3 | 2 (1) |
| MM-ToolSandbox | 5 | 5 | 1 | 1 (1) |
| Total | 34 | 34 | 8 | 8 (4) |
| Tools: audited; Mut.: in scope as mutating; Impl.: distinct implementations behind the cells; Classes: every cell, candidate included; (HL): headline-eligible. | ||||
| # | Bench. @ commit | Tool : line | Class | Tier | Quoted evidence and disposition |
| 1 | MedAgentBench @ 9926011 | POST branch, init.py:85-91 | Phantom Effect | agent-visible ( prompt_template , tool_return ) | “…executed successfully” returned with no write; payload never re-read. |
| 2 | tau2-bench @ c3398666 | refuel_data , telecom/tools.py:607-657 | Unenforced Precondition | agent-visible ( docstring ) | “must be Active” check commented out (selective disable). |
| 3 | tau2-bench @ c3398666 | cancel_reservation , airline/tools.py:315, 363-368 | Partial Effect | maintainer-annotated: not headline-eligible | “Seats release not implemented…!!!” (367); the 689 TODO of another tool asks “What about in cancel_reservation?”. Never restored; per-episode reset (Reset Leak withdrawn). |
| 4 | MedAgentBench @ 9926011 | write graders, refsol.py (SHA-256-pinned; not in repo) | Ungrounded Oracle (evaluator property, § III ) | n/a (grader source) | “POST request accepted” gates extract_posts , graded from transcript, not live state. |
| 5 | AgentDojo @ 089ed46 | update_scheduled_transaction , banking_client.py:115-151 | Ignored Argument + Phantom Effect | agent-visible ( docstring ) | recurring guarded by truthiness ( True only); returns “…updated” unconditionally. |
| 6 | AgentDojo @ 089ed46 | reserve_car_rental , travel_booking_client.py:382-400 | Ignored Argument | agent-visible ( docstring ) | end_time dropped from state, kept in success string. No task exercises this tool. |
| Benchmark : tool | Defect class | Basis | Exposed / total | Gold-call | Oracle grounding |
| tau2 airline : cancel_reservation | Partial Effect | whole_state_hash | 50 / 50 | 7 | state_grounded |
| tau2 telecom : refuel_data | Unenforced Precondition | exact_field (1120) + collection_only (15) | 1135 / 2285 | 1120 | state_grounded |
| AgentDojo banking : update_scheduled_transaction | Ignored Argument | exact_field (1) + unresolved (2) | 3 / 16 | 4 | state_grounded |
| AgentDojo travel : reserve_car_rental | Ignored Argument | exact_field (1) + unresolved (18) | 19 / 20 | 0 | mixed |
| MedAgentBench : post-write | Phantom Effect | transcript (not state) | undefined | n/a | 60 transcript / 90 mixed / 150 no-oracle / 0 state (of 300) |
| MM-ToolSandbox : venmo_social | Ignored Argument | n/c | n/c | n/c | n/c |
| Shipped PASS | Indep. PASS | State changed | Verdict changed | |||||
| Constr. | Tasks | unp. | pat. | unp. | pat. | unp. | pat. | |
| F2 C1 | 1120 | 312 | 0 | 0 | 1120 | 1120 | 0 | 312 |
| F2 C2 | 1120 | 1120 | 0 | 0 | 1120 | 1120 | 0 | 1120 |
| F3 | 6 of 7 | 0 | 0 | 0 | 6 | 6 | 0 | 0 |
| unp.: unpatched; pat.: patched; the independent verdict is unanimous in every arm, so each 2 2 is its margins. F3’s seventh task was infeasible ( Too many reservations ). | ||||||||
| Operator | n drawn | eff. n | detected | missed | untestable | error | Recall (Wilson LB) |
| M-PHANTOM | 7 | 7 | 5 | 0 | 0 | 2 | 0.359 |
| M-PRECOND | 7 | 5 | 1 | 6 | 0 | 0 | 0.020 |
| M-IGNARG | 21 | 15.13 | 4 | 12 | 1 | 4 | 0.066 |
| M-PARTIAL | 8 | 6 | 1 | 7 | 0 | 0 | 0.018 |
| M-INVAR | 8 | 8 | 0 | 8 | 0 | 0 | 0.000 |
| M-RESET | n/a | n/a | n/a | n/a | n/a | n/a | no data |