cs.CROct 6, 2026

When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback

Authors: Kavienan Jegatheesan, Gayathri Lihinikaduarachchi

Abstract

Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect and correct corrupted tool call outputs? We study this through a controlled corruption framework where a hidden interceptor replaces tool call results with plausible incorrect information on targeted problems. We evaluate agents across 31 problems under four verification designs including no verification (baseline), mandatory same-context reflection, optional fresh-context verification, and optional structural verification. Without verification, corruption causes dramatic accuracy loss, from 100% down to 72.4%. Mandatory reflection fully recovers this performance to 100%. Optional verification improves accuracy only when models actively invoke it. Our results show that checking frequency is strongly associated with robustness differences, while unequal invocation prevents a controlled comparison of verifier quality. A supporting recovery experiment shows that full problem restart succeeds in 100% of cases after explicit detection. These findings demonstrate that verifier availability and verification policy are separate components of mathematical-agent reliability. Mandatory policies enforce verification while optional policies depend on the model's own choice to invoke it.

Explore similar work

CardsList
  1. Pessimistic Verification for Open Ended Math Questions

    Nov 26, 2025Yanxing Huang, Zihan Tang, Zejin Lin +2LLM Answer VerificationLLM Mathematical Reasoning

  2. AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

    Jul 31, 2026Rui Zou, Yutao Zhu, Mengqi Wei +1LLM Self-CorrectionLLM Answer Verification

  3. Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

    Jun 24, 2026Yang Tian, Zhengpeng Shi, Yu Zhou +1AI Agent ReliabilityAI Agent Benchmarks