Tool-Augmented Language Model Agents

Latest papers 132

All topics
CardsList
  1. Retrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-Wiki

    May 25, 2026Haoliang Ming, Feifei Li, Xiaoqing Wu +1Multi-Hop QARetrieval-Augmented Generation

  2. How Many Tools Should an LLM Agent See? A Chance-Corrected Answer

    May 23, 2026Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh +1Tool-Use EvaluationTool-Augmented Language Model Agents

  3. ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

    May 22, 2026Xianzhong Ding, Yangyang Yu, Changwei Liu +1Tool-Augmented Language Model AgentsAI Coding Agents

  4. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    May 21, 2026Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11LLM Safety BenchmarksComputer-Use Agent Benchmarks

  5. SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents

    May 21, 2026Mehrdad Saberi, Keivan Rezaei, Soheil FeiziMulti-Hop QATool-Augmented Language Model Agents

  6. Terminal-World: Scaling Terminal-Agent Environments via Agent Skills

    May 20, 2026Zihao Cheng, Hongru Wang, Zeming Liu +6Tool-Augmented Language Model AgentsTerminal Agents

  7. EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL

    May 18, 2026Minrui Xu, Zilin Wang, Mengyi DENG +12Tool-Augmented Language Model AgentsAgentic RL

  8. Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

    May 17, 2026Yuxuan Lu, Ziyi Wang, Yingzhou Lu +12Tool-Augmented Language Model AgentsSynthetic Data Generation

  9. To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

    May 16, 2026Wei Shi, Ziheng Peng, Sihang Li +4Tool-Augmented Language Model AgentsCausal Interventions in Language Models

  10. R2V Agent: Teaching SLMs When to Ask for Help

    May 15, 2026Raghu Vamshi Hemadri, Humaira Firdowse Mohammed, Rishabh Maheshwary +5AI Agent ReliabilityTool-Augmented Language Model Agents

  11. WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

    May 11, 2026Shuangrui Ding, Xuanlang Dai, Long Xing +14Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  12. Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

    May 11, 2026Tz-Huan Hsu, Jheng-Hong Yang, Jimmy LinTool-Augmented Language Model AgentsAgentic Search

  13. LLM Agents Already Know When to Call Tools -- Even Without Reasoning

    May 10, 2026Chung-En Sun, Linbo Liu, Ge Yan +2Tool-Augmented Language Model AgentsAI Agent Benchmarks

  14. SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks

    May 9, 2026Jinchao Hu, Meizhi Zhong, Kehai Chen +1Tool-Augmented Language Model AgentsAgentic Search

  15. Tool Calling is Linearly Readable and Steerable in Language Models

    May 8, 2026Zekun Wu, Ze Wang, Seonglae Cho +4AI Agent ReliabilityTool-Augmented Language Model Agents

  16. Switchcraft: AI Model Router for Agentic Tool Calling

    May 8, 2026Sharad Agarwal, Pooria Namyar, Alec Wolman +3Tool-Augmented Language Model AgentsCost-Aware Inference

  17. MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

    May 7, 2026Maximillian Chen, Xuanming Zhang, Michael Peng +3Tool-Augmented Language Model AgentsMultimodal Large Language Models

  18. AgenticRAG: Agentic Retrieval for Enterprise Knowledge Bases

    May 7, 2026Susheel Suresh, Hazel Mak, Shangpo Chou +2Tool-Augmented Language Model AgentsAgentic Search

  19. Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents

    Apr 30, 2026Kaituo Zhang, Zhen Xiong, Mingyu Zhong +4Tool-Augmented Language Model AgentsAgentic Reasoning