Tool-Using Agents

Momentum

31 papers in the last four weeks, up 210% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 199

All topics
CardsList
  1. R2V Agent: Teaching SLMs When to Ask for Help

    May 15, 2026Raghu Vamshi Hemadri, Humaira Firdowse Mohammed, Rishabh Maheshwary +5AI Agent ReliabilityTool-Augmented Language Model Agents

  2. ColPackAgent: Agent-Skill-Guided Hard-Particle Monte Carlo Workflows for Colloidal Packing

    May 15, 2026Lijie Ding, Changwoo DoAgentic WorkflowsAI Agent Benchmarks

  3. Multi-Agentic Approach for History Matching of Oil Reservoirs

    May 14, 2026Linar Samigullin, Sergei Shumilin, Evgeny BurnaevMulti-Agent OrchestrationLLM Agent Orchestration

  4. CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

    May 13, 2026Mingzhi Zhu, Michele Merler, Raju Pavuluri +1Software Engineering AgentsAgentic Code Generation

  5. RS-Claw: Progressive Active Tool Exploration via Hierarchical Skill Trees for Remote Sensing Agents

    May 13, 2026Liangtian Liu, Zeyuan Wang, Ziyu Li +8Remote Sensing Image UnderstandingTool-Using Agents

  6. Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling

    May 13, 2026Coleman Hooper, Minwoo Kang, Suhong Moon +7Tool-Using AgentsLLM Agents

  7. No Action Without a NOD: A Heterogeneous Multi-Agent Architecture for Reliable Service Agents

    May 12, 2026Zixu Yang, Hang Zheng, Nan Jiang +5AI Agent ReliabilityMulti-Agent LLM Systems

  8. When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents

    May 12, 2026Xiaolin Zhou, Aojie Yuan, Zheng Luo +12Domain RandomizationAI Agent Benchmarks

  9. Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations

    May 12, 2026Junjue Wang, Weihao Xuan, Heli Qi +7LLM Agent EvaluationGeospatial Reasoning

  10. ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

    May 11, 2026Yuanyang Li, Xue Yang, Longyue Wang +2LLM Agent EvaluationAI Agent Benchmarks

  11. AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care

    May 8, 2026Preyash Yadav, Michelle Cohn, Priyanka Koppolu +5Alzheimer's DiseaseTool-Using Agents

  12. Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling

    May 8, 2026Naoki Otani, Nikita Bhutani, Hannah Kim +2Long-Horizon PlanningTool-Using Agents

  13. AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents

    May 8, 2026Zhengkang Guo, Yiyang Li, Lin Qiu +7LLM Agent EvaluationLong-Horizon Agent Evaluation

  14. MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents

    May 7, 2026Ashwani Anand, Ivi Chatzi, Ritam Raha +1Computer-Use Agent BenchmarksAI Agent Reliability

  15. LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents

    May 6, 2026Yijun Lu, Rui Ye, Yuwen Du +3Web Search AgentsTool-Using Agents

  16. Beyond State Machines: Executing Network Procedures with Agentic Tool-Calling Sequences

    May 4, 2026Purna Sai Garigipati, Onur Ayan, Kishor Chandra Joshi +1LLM Agent OrchestrationTool-Using Agents

  17. Bidirectional Semantic Complementary Tool Retrieval for Remote Sensing Agents

    Apr 29, 2026Zeyuan Wang, Dongyang Hou, Cheng Yang +10Remote SensingTool-Using Agents

  18. When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors

    Apr 23, 2026Chenghao Yang, Yuning Zhang, Zhoufutu Wen +4Tool-Augmented Language Model AgentsLLM Agent Evaluation

  19. ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

    Apr 20, 2026Xirui Li, Ming Li, Ion Stoica +2Agent EvaluationAI Agent Benchmarks