Tool-Using Agents

Momentum

31 papers in the last four weeks, up 210% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 199

All topics
CardsList
  1. DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration

    Oct 8, 2026Yingda Shen, Yuxiang Wang, Kunyu Feng +7LLM Agent HarnessesAgent Harness Optimization

  2. TaReD: Tool-Aware Recursive Decomposition for Long-Horizon Tasks

    Oct 8, 2026Wei-Xiang Mao, Zhi-Kai Chen, De-Chuan Zhan +1Tool-Using AgentsLong-Horizon Agent Tasks

  3. Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents

    Oct 7, 2026Obada KraishanTool-Using AgentsLLM Agent Reliability

  4. ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

    Oct 6, 2026Arkajyoti Chakraborty, Aryan Tayal, Ishika Agarwal +5Tool-Using AgentsAI Agent Benchmarks

  5. M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding

    Oct 6, 2026Jinsong Zhang, Kejun Wu, Ming Zhu +3Monocular Depth EstimationTool-Using Agents

  6. From Evidence to Action: How Tool-Using Agents Fail

    Oct 6, 2026Hongzhan Lin, Shidong Cao, Ziyang Luo +3Agent Failure AnalysisEvidence-Grounded Reasoning

  7. Programmatic Search Agents: Extending Agentic Search Beyond Query Reformulation

    Oct 5, 2026Jiaming Qian, Huiyan Yang, Mandi Liu +7Agentic SearchWeb Search Agents

  8. Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

    Oct 5, 2026Chubin Zhang, Zhenglin Wan, Xingrui Yu +4LLM Agent EvaluationTool-Using Agents

  9. Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents

    Sep 30, 2026Gabriel TuriniciLLM Agent Skill LearningTool-Using Agents

  10. Growing an Agent/Prover Interface: Evolutionary Tool Design for Cost-Efficient Theorem Proving in Rocq and Lean

    Sep 30, 2026Jules Viennot, Guillaume Baudart, Marc LelargeAutomated Theorem ProvingEvolutionary Optimization

  11. Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools

    Sep 30, 2026Shijia Ge, Alex Zhou, Jianshu Zeng +12Robotic Tool UseTool-Using Agents

  12. Encore: Few-Shot Agentic Discovery of Manipulation Strategies

    Sep 29, 2026Yifan Kang, Zihan Wang, Zhiwen Fan +1Robot Policy LearningLanguage-Conditioned Robot Manipulation

  13. Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

    Sep 29, 2026Ido Levy, Asaf Yehudai, Segev Shlomov +2Proactive AssistanceTool-Using Agents

  14. The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents

    Sep 29, 2026Xueqi Li, Jingjie Ning, Yibo KongLLM Agent EvaluationTool-Using Agents

  15. Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models

    Sep 28, 2026Junru Zhu, Shiming Xie, Aime Lu Fan Chen +4Agent Failure AnalysisAI Agent Benchmarks

  16. SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents

    Sep 28, 2026Xinjie Shen, Junran Wang, Rongzhe Wei +1AI Agent SecurityAI Agent Security Benchmarks

  17. Commitment Hierarchies under Intent Revision: A Belief-Revision Account of Salvage in Tool-Use Agents

    Sep 28, 2026Spandan Ghose ChowdhuryBelief RevisionLLM Planning

  18. Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

    Sep 24, 2026Zexi Li, Yehang Zhang, Wenqian Li +12Language-Conditioned Robot ManipulationTool-Using Agents

  19. UniDataAgent: An Ontology-Grounded Agent for Enterprise Question-to-Report Automation

    Sep 23, 2026Yutai Duan, Yahui Zhao, Zhangti Li +5Tool-Using AgentsLLM Agents

  20. Canonical Procedural Actions: An Auditable Annotation Protocol for Tool-Use Agent Traces

    Sep 21, 2026Songqi Li, Dongqing Li, Zheqiao ChengTool-Using Agents

  21. MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents

    Sep 21, 2026Demetris Paschalides, Moysis Symeonides, George Pallis +1AI Agent BenchmarksTool-Using Agents

  22. An Empirical Study of Harness Design for Coding Agents

    Sep 17, 2026Run-Ze Fan, Zihao Zhang, Simin Ma +6Coding AgentsAgent Harness Optimization

  23. Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

    Sep 16, 2026Yipeng Liu, Yingqiang Zhang, Feifei Li +1LLM Inference EfficiencyKV Caching

  24. TuiML: Machine Learning for AI Agents

    Sep 16, 2026Nilesh Verma, Nick Lim, Albert Bifet +1Tool-Using AgentsML Reproducibility

  25. Semantic-TVM: Structure-Preserving Trustworthy Virtual Memory for Memory-Augmented and Tool-Using Agents

    Sep 14, 2026Yu Li, Qikun Cai, Tao Huang +1Agent MemoryPrivacy-Preserving Language Models

  26. Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents

    Sep 14, 2026Zhutao Lv, Chenhao Dang, Yi Feng +5Remote SensingAI Agent Benchmarks

  27. ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents

    Sep 14, 2026Bowen Guan, Zhentao Yin, Yanming ShenLLM Agent EvaluationAI Agent Benchmarks

  28. Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

    Sep 13, 2026Arham Sethi, Arsen Kenzhebayev, Saanvi Paturi +3Tool-Augmented Language Model AgentsDeception in Language Models

  29. The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents

    Sep 10, 2026Bo Yan, Weikai Lin, Song WangTool-Using AgentsLLM Tool Use

  30. From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents

    Sep 7, 2026Yongjian Lyu, Yang Ren, Ruofei Lai +1Tool-Using AgentsRuntime Enforcement for AI Agents

  31. Speculative Macro Commit for Faster Tool-Using Agents

    Sep 3, 2026Zeyu Liu, Souvik Kundu, Peter A. BeerelTool-Augmented Language Model AgentsLLM Agent Workflow Optimization

  32. Source-Dependent Deference in Medical Imaging Agents Under Falsified Findings: A Pilot Audit

    Aug 30, 2026Ridam Roy, Md Shahriar Rashid, Md. Rajib MiaTool-Use EvaluationMedical VQA

  33. Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses

    Aug 24, 2026Jiachen Xu, Torben Bach Pedersen, Zhongming Yao +2Agent Failure AnalysisAI Agent Reliability

  34. One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

    Aug 20, 2026Zhuochun Li, Youngmin Ko, Ali Keramati +11Business Process AutomationAI Agent Benchmarks

  35. PIPES: Securing Agent Perception with Provenance and Priors

    Aug 13, 2026Sanjay Kariyappa, Severin Klingler, G. Edward SuhData ProvenanceAI Agent Security

  36. FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows

    Aug 10, 2026Shuo Hao, You Lu, Bihuan Chen +1Agentic WorkflowsLLM Agent Workflow Optimization

  37. DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents

    Aug 10, 2026You Lu, Kun Zhang, Bihuan Chen +1Tool-Augmented Language Model AgentsTool-Using Agents

  38. OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

    Aug 9, 2026Andrea Caciolai, Pere-Lluís Huguet Cabot, Chierh Cheng +11AI Agent EvaluationAI Agent Benchmarks

  39. The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task

    Aug 9, 2026Marc Alier Forment, María José Casañ Guerrero, Francisco José García-Peñalvo +1AI Coding AgentsLLM Agent Evaluation

  40. WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance

    Aug 7, 2026Zhi Li, Tao Zhou, Yeqing Li +2LLM Agent EvaluationTool-Using Agents

  41. When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment

    Aug 6, 2026Weihong Lin, Lin Sun, Xiangzheng ZhangLLM Agent EvaluationTool-Using Agents

  42. ETA: A New Agentic Paradigm for Embodied Tasks

    Aug 4, 2026Yitong Chen, Zezheng Huai, Sixian Li +7Tool-Using AgentsRobot Task Planning

  43. SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite Intelligence

    Aug 4, 2026Longji He, Jeto XuEdge ComputingLLM Agent Orchestration

  44. WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

    Aug 4, 2026Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang +3Multi-Agent CollaborationAI Agent Security Benchmarks

  45. Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

    Aug 4, 2026Can Wang, Haoran Chen, Li Yu +4Tool-Using AgentsLLM Tool Use

  46. ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

    Aug 3, 2026Vernon Toh, Navonil Majumder, Zhengyuan Liu +2LLM Agent EvaluationAI Agent Benchmarks

  47. Homebot: A Personal AI Agent for Conversational Home Assistance and Automation

    Aug 3, 2026Shengyuan Ye, Yixin Zhang, Han Liang +3Tool-Augmented Language Model AgentsTool-Using Agents