cs.AIOct 7, 2026

Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents

Authors: Obada Kraishan

Organizations: College of Media and Communication Texas Tech University Lubbock, TX, USA

Abstract

Tool-using agents are usually scored on whether they finish a task while the tools work. Deployments are less forgiving: services time out, endpoints disappear, parameter names change, and results come back well formed but wrong. Prior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents. Wrapping the executable environments of an established function-calling benchmark in a fault-injection layer, we inject one of four typed faults at a controlled point in the trajectory and record whether the agent notices, changes plan, recovers the task, or repeats itself. Six models from three families, half of them reasoning variants, ran 1,920 trials over 24 multi-step tasks. Agents treat a failure as a problem in 91.3% of trials when the tool returns an explicit error, but in 58.8% of trials when it returns a plausible wrong value, against a 26.8% rate of reporting problems when nothing was wrong. Reasoning models are not better placed: paired against instruct siblings, they notice less (-9.3 points, p < .001) and change plan more (+10.4 points, p < .001), and recovery is unchanged (p = .512). Because agents are stochastic, two fault-free runs of the same task end in the same state only 63.3% of the time; against that baseline, only a missing tool clearly lowers recovery (39.9%), while timeouts, schema drift, and corruption stay within run-to-run variation. After a fault, agents return to the same tool three or more times in a row in up to 22.2% of trials, though strictly identical repeats are rare. A prompt line asking the agent to check each result did not move detection. Agents respond to the error channel rather than to the content of what a tool returns, so failures that stay inside the expected format pass through.

Figures & tables

Explore similar work

CardsList
  1. Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models

    Sep 28, 2026Junru Zhu, Shiming Xie, Aime Lu Fan Chen +4Latent Failure Patterns

  2. Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

    Sep 13, 2026Arham Sethi, Arsen Kenzhebayev, Saanvi Paturi +3HonestySystemverilog Assertions

  3. When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

    Jun 4, 2026Dongsheng Zhu, Xuchen Ma, Yucheng Shen +5Large Language Model AgentsFault Tolerance