Tool-Augmented Language Model Agents

Latest papers 132

All topics
CardsList
  1. LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

    Jun 18, 2026Md Nayem Uddin, Amir Saeidi, Eduardo Blanco +1Tool-Augmented Language Model AgentsCustomer Support Automation

  2. Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns

    Jun 16, 2026Shiqi He, Yue Cui, Feijie Wu +5Tool-Augmented Language Model AgentsLLM Agent Skill Learning

  3. MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

    Jun 15, 2026Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh +1Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  4. ACCORD: Action-Conditioned Contextual Grounding for Language Agents

    Jun 15, 2026Lai Jiang, Cheng Qian, Zhenhailong Wang +3LLM GroundingTool-Augmented Language Model Agents

  5. ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents

    Jun 13, 2026Rahul Suresh Babu, Laxmipriya Ganesh IyerAI Agent ReliabilityTool-Use Evaluation

  6. EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

    Jun 13, 2026Siyuan Zhang, Jian Zong, Junyu Wang +7Audio QATool-Augmented Language Model Agents

  7. HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents

    Jun 11, 2026Yaxin Du, Yifan Zhou, Yujie Ge +7Tool-Augmented Language Model AgentsLLM Agent Workflow Optimization

  8. Recursive Agent Harnesses

    Jun 11, 2026Elias Lumer, Sahil Sen, Kevin Paul +1Tool-Augmented Language Model AgentsLLM Agent Orchestration

  9. QueryWeaver: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph Generation

    Jun 6, 2026Aishwarya Chakravarthy, Vidhi Kulkarni, Duen Horng ChauTool-Augmented Language Model AgentsTool-Use Planning

  10. ToolChoiceConfusion: Causal Minimal Tool Filtering for Reliable LLM Agents

    Jun 4, 2026Rahul Suresh Babu, Laxmipriya Ganesh IyerAI Agent ReliabilityTool-Use Evaluation

  11. Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

    Jun 4, 2026Christopher J. Wedge, Joshua Stutter, Danny Dixon +1Multi-Hop QATool-Augmented Language Model Agents

  12. Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

    Jun 3, 2026Haoyu Sun, Wenxuan Wang, Mingyang Song +5Tool-Augmented Language Model AgentsLLM Agent Evaluation

  13. ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents

    Jun 2, 2026Anjie Liu, Yan Song, Zhixun Chen +3Tool Use in VLMsTool-Augmented Language Model Agents

  14. WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

    Jun 1, 2026Hengrui Gu, Xiaotian Han, Kaixiong ZhouTool-Augmented Language Model AgentsTool-Using Agents

  15. SkillsInjector: Dynamic Skill Context Construction for LLM Agents

    May 28, 2026Yanchao Li, Wanhao Liu, Ben Gao +5Tool-Augmented Language Model AgentsLLM Agent Skill Learning

  16. Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

    May 28, 2026Lorenz Kutschka, Bernhard GeigerTool-Augmented Language Model AgentsAI Agent Evaluation

  17. VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data

    May 28, 2026Di Zhu, Yu Yvonne Wu, Hong Jia +3Tool-Augmented Language Model AgentsLLM Agents

  18. GrepSeek: Training Search Agents for Direct Corpus Interaction

    May 28, 2026Alireza Salemi, Chang Zeng, Atharva Nijasure +4Tool-Augmented Language Model AgentsAgentic Search

  19. A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

    May 27, 2026Tomer Keren, Nitay Calderon, Asaf Yehudai +3Computer-Use Agent BenchmarksTool-Use Evaluation

  20. Plan Before Search: Search Agents Need Plan

    May 27, 2026Zhipeng Qian, Zihan Liang, Yufei Ma +7Multi-Hop QATool-Augmented Language Model Agents

  21. Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents

    May 27, 2026Xinze Li, Yuhang Zang, Yixin Cao +1Tool-Augmented Language Model AgentsLLM Agent Skill Retrieval

  22. Retrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-Wiki

    May 25, 2026Haoliang Ming, Feifei Li, Xiaoqing Wu +1Multi-Hop QARetrieval-Augmented Generation

  23. How Many Tools Should an LLM Agent See? A Chance-Corrected Answer

    May 23, 2026Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh +1Tool-Use EvaluationTool-Augmented Language Model Agents

  24. ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

    May 22, 2026Xianzhong Ding, Yangyang Yu, Changwei Liu +1Tool-Augmented Language Model AgentsAI Coding Agents

  25. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    May 21, 2026Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11LLM Safety BenchmarksComputer-Use Agent Benchmarks

  26. SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents

    May 21, 2026Mehrdad Saberi, Keivan Rezaei, Soheil FeiziMulti-Hop QATool-Augmented Language Model Agents

  27. Terminal-World: Scaling Terminal-Agent Environments via Agent Skills

    May 20, 2026Zihao Cheng, Hongru Wang, Zeming Liu +6Tool-Augmented Language Model AgentsTerminal Agents

  28. EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL

    May 18, 2026Minrui Xu, Zilin Wang, Mengyi DENG +12Tool-Augmented Language Model AgentsAgentic RL

  29. Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

    May 17, 2026Yuxuan Lu, Ziyi Wang, Yingzhou Lu +12Tool-Augmented Language Model AgentsSynthetic Data Generation

  30. To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

    May 16, 2026Wei Shi, Ziheng Peng, Sihang Li +4Tool-Augmented Language Model AgentsCausal Interventions in Language Models

  31. R2V Agent: Teaching SLMs When to Ask for Help

    May 15, 2026Raghu Vamshi Hemadri, Humaira Firdowse Mohammed, Rishabh Maheshwary +5AI Agent ReliabilityTool-Augmented Language Model Agents

  32. WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

    May 11, 2026Shuangrui Ding, Xuanlang Dai, Long Xing +14Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  33. Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

    May 11, 2026Tz-Huan Hsu, Jheng-Hong Yang, Jimmy LinTool-Augmented Language Model AgentsAgentic Search

  34. LLM Agents Already Know When to Call Tools -- Even Without Reasoning

    May 10, 2026Chung-En Sun, Linbo Liu, Ge Yan +2Tool-Augmented Language Model AgentsAI Agent Benchmarks

  35. SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks

    May 9, 2026Jinchao Hu, Meizhi Zhong, Kehai Chen +1Tool-Augmented Language Model AgentsAgentic Search

  36. Tool Calling is Linearly Readable and Steerable in Language Models

    May 8, 2026Zekun Wu, Ze Wang, Seonglae Cho +4AI Agent ReliabilityTool-Augmented Language Model Agents

  37. Switchcraft: AI Model Router for Agentic Tool Calling

    May 8, 2026Sharad Agarwal, Pooria Namyar, Alec Wolman +3Tool-Augmented Language Model AgentsCost-Aware Inference

  38. MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

    May 7, 2026Maximillian Chen, Xuanming Zhang, Michael Peng +3Tool-Augmented Language Model AgentsMultimodal Large Language Models

  39. AgenticRAG: Agentic Retrieval for Enterprise Knowledge Bases

    May 7, 2026Susheel Suresh, Hazel Mak, Shangpo Chou +2Tool-Augmented Language Model AgentsAgentic Search

  40. Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents

    Apr 30, 2026Kaituo Zhang, Zhen Xiong, Mingyu Zhong +4Tool-Augmented Language Model AgentsAgentic Reasoning