Tool-Using Agents

Momentum

31 papers in the last four weeks, up 210% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 199

All topics
CardsList
  1. Declarative Skills for AI Agents in Knowledge-Grounded Tool-Use Workflows

    Jun 5, 2026M. Danish Lim, I. Danial Bin Sharudin, Wen Han Chen +2LLM Agent OrchestrationLLM Agent Workflow Optimization

  2. Archi: Agentic Operations at the CMS Experiment

    Jun 3, 2026Pietro Lugato, Luca Lavezzo, Jason Mohoney +16Retrieval-Augmented GenerationTool-Using Agents

  3. Uncertainty-Aware Clarification in LLM Agents with Information Gain

    Jun 2, 2026Mengyi Deng, Zhiwei Li, Xin Li +4Clarification Question GenerationTool-Using Agents

  4. WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

    Jun 1, 2026Hengrui Gu, Xiaotian Han, Kaixiong ZhouTool-Augmented Language Model AgentsTool-Using Agents

  5. Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools

    Jun 1, 2026Bardia Mohammadi, Lars Klein, Akhil Arora +1Data LeakageLLM Agent Security

  6. MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation

    Jun 1, 2026Wenhao Wang, Peizhi Niu, Gongyi Zou +9LLM Agent EvaluationAI Agent Benchmarks

  7. Beyond One-shot: AI Agents for Learning in Field Experiments

    Jun 1, 2026Junjie Luo, Ritu Agarwal, Gordon GaoTool-Using AgentsAutonomous Experimentation

  8. SkillSmith: Co-Evolving Skills and Tools for Self-Improving Agent Systems

    May 31, 2026Yangbo Wei, Zhen Huang, Shaoqiang Lu +4Tool-Using AgentsAgent Skill Learning

  9. CAREAgent: Clinical Agent with Structured Reasoning and Tool-Integrated for Order Generation

    May 31, 2026Ruihui Hou, Ziyue Huai, Chennuo Zhang +5Tool-Using AgentsClinical Reasoning

  10. Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning

    May 30, 2026Chishui Chen, Jiaye Lin, Te Sun +6LLM Agent Skill LearningTool-Using Agents

  11. A Multi-AI-agent Framework Enabling End-to-end Finite Element Analysis for Solid Mechanics Problems

    May 28, 2026Titu Ranjan Sarker, Muhammed Jawaad Zulqernine, Ling Yue +3Multi-Agent LLM SystemsStructural Mechanics

  12. Scaling Laws for Agent Harnesses via Effective Feedback Compute

    May 28, 2026Xuanliang Zhang, Dingzirui Wang, Keyan Xu +2Inference-Time ScalingAgent Harness Optimization

  13. AIRGuard: Guarding Agent Actions with Runtime Authority Control

    May 27, 2026Suliu Qin, Haomin Zhuang, Yujun Zhou +2AI Agent SecurityTool-Using Agents

  14. AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

    May 27, 2026Kou Shi, Ziao Zhang, Shiting Huang +7Tool-Use EvaluationAI Agent Benchmarks

  15. ORCA: An End-to-End Interactive Copilot for Optimized Root Cause Analysis

    May 26, 2026Phi Nguyen Xuan, Nicholas Tagliapietra, Lavdim Halilaj +2AI-Assisted Decision MakingCausal Discovery

  16. MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

    May 24, 2026Xuanye Zhang, Yongsen Zheng, Zhuqin Xu +5LLM Agent SecurityTool-Using Agents

  17. Agent Manufacturing: Foundation-Model Agents as First-Class Industrial Entities

    May 24, 2026Yilei ZhangManufacturingTool-Using Agents

  18. How Many Tools Should an LLM Agent See? A Chance-Corrected Answer

    May 23, 2026Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh +1Tool-Use EvaluationTool-Augmented Language Model Agents

  19. Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning

    May 21, 2026Banghao Chi, Yining Xie, Mingyuan Wu +9Tool-Using AgentsLLM Agent Training

  20. TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks

    May 21, 2026Zhaoyang Chu, Jiarui Hu, Xingyu Jiang +8Computer-Use Agent BenchmarksTerminal Agents

  21. When Skills Don't Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity

    May 19, 2026Samuel Jacob Chacko, James Hugglestone, Chashi Mahiul Islam +1LLM Agent EvaluationLLM Agent Skill Learning

  22. Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

    May 18, 2026Rishi Jha, Harold Triedman, Arkaprabha Bhattacharya +1Agent Failure AnalysisAI Agent Reliability

  23. TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents

    May 16, 2026Zhiqiang Liu, Wenhui Dong, Yilang Tan +3Tool-Use EvaluationAI Agent Benchmarks

  24. CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

    May 15, 2026Haolin Chen, Deon Metelski, Leon Qi +30HealthcareLong-Horizon Agent Evaluation