Tool-Use Evaluation

Momentum

12 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 67

All topics
CardsList
  1. GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

    Apr 17, 2026Jize Wang, Xuanxuan Liu, Yining Li +7Tool-Use EvaluationLong-Horizon Agent Evaluation

  2. The Tool Illusion: Rethinking Tool Use in Web Agents

    Apr 3, 2026Renze Lou, Baolin Peng, Wenlin Yao +5Tool-Use EvaluationWeb Agents

  3. Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

    Apr 1, 2026Alibek Kaliyev, Artem MaryanskyyAI Agent ReliabilityTool-Use Evaluation

  4. CCTU: A Benchmark for Tool Use under Complex Constraints

    Mar 16, 2026Junjie Ye, Guoqiang Zhang, Wenjie Fu +4Tool-Use EvaluationAI Agent Benchmarks

  5. When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

    Oct 22, 2025Tanya Shourya, Yingfan Wang, Zhaoyi Joey Hou +3Tool-Use EvaluationAI Agent Benchmarks