Tool-Using Agents

Momentum

31 papers in the last four weeks, up 210% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 199

All topics
CardsList
  1. When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations

    Jul 1, 2026Qian Chen, Chengyuan Liu, Xin YuCustomer Support AutomationTool-Using Agents

  2. Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

    Jul 1, 2026Song-Lin Lv, Weiming Wu, Rui Zhu +2Fine-TuningDistribution Shift

  3. Entity Binding Failures in Tool-Augmented Agents

    Jun 29, 2026Rahul Suresh Babu, Shashank IndukuriTool-Using AgentsLLM Agent Safety

  4. How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

    Jun 24, 2026David Akinpelu, Akintonde Abbas, Rereloluwa Alimi +1LLM Agent EvaluationAI Agent Benchmarks

  5. Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

    Jun 24, 2026Yang Tian, Zhengpeng Shi, Yu Zhou +1AI Agent ReliabilityAI Agent Benchmarks

  6. SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis

    Jun 23, 2026Yucheng Yuan, Yuanfeng Ji, Zhongxiao Li +1AI Agent BenchmarksTool-Using Agents

  7. Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning

    Jun 22, 2026Jiaqiang TangTool-Using AgentsSelf-Improving Agents

  8. Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents

    Jun 19, 2026Chubin Zhang, Zhenglin Wan, Xingrui Yu +5AI Agent ReliabilityTool-Use Evaluation

  9. AgentMeter: Evaluating Model-CLI Matching for CLI-Based Local Task-Solving Agents

    Jun 19, 2026Han Chi, Jiaxin Qi, Yan Cui +2LLM Agent EvaluationAI Agent Benchmarks

  10. LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

    Jun 18, 2026Md Nayem Uddin, Amir Saeidi, Eduardo Blanco +1Tool-Augmented Language Model AgentsCustomer Support Automation

  11. ToolChain-CRC: Conformal Risk Control for Agentic AI Under Retrieval and Tool-Use Drift

    Jun 16, 2026Jeffery Opoku, David BanaheneAI Agent ReliabilityDistribution Shift

  12. A T-API-Compliant ReAct Agentic Loop for Optical Networks: Generic vs. Domain-Specific Tool Abstractions

    Jun 16, 2026Seyed Morteza Ahmadian, Paolo Monti, Carlos NatalinoTool-Using AgentsAgentic Reasoning

  13. GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents

    Jun 15, 2026Rahul Suresh Babu, Rohit ShuklaClarification Question GenerationTool-Using Agents

  14. Can Agents Infer Environment from Interaction? Evidence from Agentic Automata Learning

    Jun 15, 2026Reef Menaged, Gili Lior, Shauli Ravfogel +2LLM Agent EvaluationActive Learning

  15. An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios

    Jun 15, 2026Hankyul Baek, Jaewon Noh, Sang Seo +5Data LeakageLLM Agent Evaluation

  16. Not All Skills Help: Measuring and Repairing Agent Knowledge

    Jun 13, 2026Yixuan Wang, Yiyang Zhou, Yiming Liang +4Causal AttributionLLM Agent Skill Learning

  17. PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions

    Jun 12, 2026Chenxin Li, Zhengyao Fang, Zhengyang Tang +18Mobile GUI AutomationTool-Using Agents

  18. Strategic Decision Support for AI Agents

    Jun 10, 2026Shayan Kiyani, Sima Noorani, George Pappas +1AI Agent ReliabilityTool-Using Agents

  19. Agents All the Way Down; A Methodology for Building Custom AI Agents from Substrate to Production

    Jun 10, 2026Marc Alier Forment, Juanan Pereira, Francisco José García-Peñalvo +1LLM Agent OrchestrationTool-Using Agents

  20. MedCTA: A Benchmark for Clinical Tool Agents

    Jun 10, 2026Tajamul Ashraf, Hyewon Jeong, Fida Mohammad Thoker +1Agent EvaluationAI Agent Benchmarks

  21. ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories

    Jun 9, 2026Siyuan Luo, Nairong Zheng, Lin Zhou +6Synthetic Data GenerationTool-Using Agents

  22. T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

    Jun 9, 2026Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin +12Customer Support AutomationLLM Agent Evaluation

  23. Anything2Skill: Compiling External Knowledge into Reusable Skills for Agents

    Jun 8, 2026Qianjun Pan, Yutao Yang, Junsong Li +5LLM Agent MemoryLLM Agent Skill Learning

  24. Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps

    Jun 8, 2026Xiaofeng Lin, Yukai Yang, Daniel Guo +3Jailbreak AttacksLLM Agent Security

  25. Projecting the Emerging Mindset of SWE Agent by Launching a Wild Code Understanding Journey

    Jun 7, 2026Zhengyi Zhuo, Yan LiuSoftware EngineeringSoftware Engineering Agents

  26. MemToolAgent: Leveraging Memory for Tool Using Agents Based on Environment and User Feedback

    Jun 6, 2026Suleyman Armagan Er, Danilo Ribeiro, Yogesh Virkar +5Tool-Using AgentsLLM Tool Use

  27. Contract2Tool: Learning Preconditions and Effects for Reliable Tool-Augmented LLM Agents

    Jun 5, 2026Rahul Suresh Babu, Laxmipriya Ganesh IyerAI Agent ReliabilityTool-Use Evaluation