Tool-Augmented Language Model Agents

Latest papers 132

All topics
CardsList
  1. Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

    Oct 1, 2026Ruiyang Si, Jianxin Bi, Shunyu Yang +9Robotic ControlTool-Augmented Language Model Agents

  2. Finding the Right Fit: Model-Harness Interactions across Agent Tasks

    Oct 1, 2026Yixuan Li, Yiyun Zhou, Yao Long Teng +6Tool-Augmented Language Model AgentsLLM Agent Evaluation

  3. TMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving

    Sep 28, 2026Shengqin Wang, Jie Jin, Yu Cheng +4Molecular OptimizationTool-Augmented Language Model Agents

  4. AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

    Sep 28, 2026Chanhee Park, Jeongho Yoon, Sungbin Han +2Multi-Hop QATool-Augmented Language Model Agents

  5. Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise

    Sep 24, 2026Arian Abbasi, Alan Aqrawi, Ted KwartlerTool-Augmented Language Model AgentsCoding Agents

  6. SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

    Sep 23, 2026Zhilong Ge, Yuting Shao, Yutao Yang +6Tool-Augmented Language Model AgentsLLM Agent Skill Learning

  7. Toolcompass: Guiding Tool Trialing, Not Suppressing It

    Sep 22, 2026Junlin Fang, Chong Zhang, Do Nguyen-Thanh +3Representation LearningTool-Augmented Language Model Agents

  8. Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

    Sep 21, 2026Lujia Bao, Qian Chen, Luyao Cheng +15Tool-Augmented Language Model AgentsSpeech Foundation Models

  9. Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

    Sep 16, 2026Mahsa Amani, Seungeon Lee, Abhisek Dash +9Tool-Augmented Language Model AgentsAgentic Search

  10. Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

    Sep 14, 2026Zixiang Chen, Sufeng Niu, Yingchi Liu +22Tool-Augmented Language Model AgentsLanguage Modeling

  11. Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

    Sep 13, 2026Arham Sethi, Arsen Kenzhebayev, Saanvi Paturi +3Tool-Augmented Language Model AgentsDeception in Language Models

  12. Speculative Macro Commit for Faster Tool-Using Agents

    Sep 3, 2026Zeyu Liu, Souvik Kundu, Peter A. BeerelTool-Augmented Language Model AgentsLLM Agent Workflow Optimization

  13. Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives

    Sep 1, 2026Haibo Jin, Suijin Wang, Xucheng Yu +2Tool-Augmented Language Model AgentsLLM Agent Orchestration

  14. TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

    Aug 24, 2026Wenhao Wu, Menghao Zhang, Xin Wang +3Tool-Augmented Language Model AgentsLLM Agent Orchestration

  15. Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

    Aug 10, 2026Shulin Tian, Ziqi Huang, Fan Zhang +3Tool-Augmented Language Model AgentsAI Agent Evaluation

  16. UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

    Aug 10, 2026Xuexiong Yin, Zechuan Chen, Yongsen Zheng +5Tool-Use EvaluationTool-Augmented Language Model Agents

  17. DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents

    Aug 10, 2026You Lu, Kun Zhang, Bihuan Chen +1Tool-Augmented Language Model AgentsTool-Using Agents

  18. Persistent Semantic Entities in Tool-Augmented LLM Systems

    Aug 8, 2026Zhaohui WangAgent MemoryTool-Augmented Language Model Agents

  19. NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

    Aug 7, 2026Aditya Katkar, Om Karkele, Kartik Mandhane +2Tool-Augmented Language Model AgentsLLM Guardrails

  20. The Bitter Lesson of Tool Calling

    Aug 6, 2026Ishan Patel, Sahil Sen, Elias Lumer +1LLM EvaluationTool-Augmented Language Model Agents

  21. PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

    Aug 3, 2026Abdulrahman AlRabah, Xiaocheng Yang, Dilek Hakkani-Tür +1AI Agent ReliabilityAI-Assisted Decision Making

  22. Homebot: A Personal AI Agent for Conversational Home Assistance and Automation

    Aug 3, 2026Shengyuan Ye, Yixin Zhang, Han Liang +3Tool-Augmented Language Model AgentsTool-Using Agents

  23. Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

    Aug 3, 2026Yuwen Wang, Tian-Hao Zhang, Minghao Cai +7Tool-Augmented Language Model AgentsAgent Evaluation

  24. SearchMaster: Grounded and Regulated Self-Play for Search Agents

    Aug 3, 2026Wentao Tan, Qiong Cao, Jiaqi Wang +1Multi-Hop QATool-Augmented Language Model Agents

  25. Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

    Jul 28, 2026Xinyi Hong, Pinjun Dong, Xinyang Yu +1Tool-Augmented Language Model AgentsLLM Tool Use

  26. E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

    Jul 26, 2026Weihuang Zheng, Tianyuan Zou, Eileen Ye +5Computer-Use Agent BenchmarksTool-Use Evaluation

  27. AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

    Jul 25, 2026Hao Jiang, Gangtao Xin, Yingdi Huang +35Tool-Augmented Language Model AgentsAgentic RL

  28. SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

    Jul 23, 2026Hao Zhang, Yiwen Zhao, Yixuan Zhang +2Tool-Augmented Language Model AgentsAudio Generation

  29. TerraLogic: A Benchmark for Hierarchical Geospatial Reasoning in Earth Observation

    Jul 14, 2026Yuhang Yan, Linchao Mou, Bokang Yang +1Spatial Reasoning BenchmarksHierarchical Planning

  30. ToFu: A White-Box, Token-Efficient Agent Harness for Researchers

    Jul 13, 2026Junhao Ruan, Yuan Ge, Bei Li +7Tool-Augmented Language Model AgentsSoftware Engineering Agents

  31. BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

    Jul 13, 2026Yuzhe Guo, Mengzhou Wu, Yuan Cao +4Tool-Augmented Language Model AgentsLanguage Model Generation Evaluation

  32. AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

    Jun 30, 2026Zhaojian Yu, Penghao Yin, Shuzheng Gao +3Tool-Augmented Language Model AgentsLLM Tool Use

  33. SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search

    Jun 30, 2026Ming Dai, Zhihong Lu, Jinjie Gu +5Tool-Augmented Language Model AgentsAgentic Search

  34. LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via State Proprioception

    Jun 29, 2026Binyan Xu, Haitao Li, Kehuan ZhangTool-Augmented Language Model AgentsLLM Agent Memory

  35. MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

    Jun 25, 2026Inderjeet Singh, Andrés Murillo, Motoyoshi Sekiya +2Multimodal RAGTool-Augmented Language Model Agents