Tool-Augmented Language Model Agents

Latest papers 132

All topics
CardsList
  1. Persistent Semantic Entities in Tool-Augmented LLM Systems

    Aug 8, 2026Zhaohui WangAgent MemoryTool-Augmented Language Model Agents

  2. NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

    Aug 7, 2026Aditya Katkar, Om Karkele, Kartik Mandhane +2Tool-Augmented Language Model AgentsLLM Guardrails

  3. The Bitter Lesson of Tool Calling

    Aug 6, 2026Ishan Patel, Sahil Sen, Elias Lumer +1LLM EvaluationTool-Augmented Language Model Agents

  4. PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

    Aug 3, 2026Abdulrahman AlRabah, Xiaocheng Yang, Dilek Hakkani-Tür +1AI Agent ReliabilityAI-Assisted Decision Making

  5. Homebot: A Personal AI Agent for Conversational Home Assistance and Automation

    Aug 3, 2026Shengyuan Ye, Yixin Zhang, Han Liang +3Tool-Augmented Language Model AgentsTool-Using Agents

  6. Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

    Aug 3, 2026Yuwen Wang, Tian-Hao Zhang, Minghao Cai +7Tool-Augmented Language Model AgentsAgent Evaluation

  7. SearchMaster: Grounded and Regulated Self-Play for Search Agents

    Aug 3, 2026Wentao Tan, Qiong Cao, Jiaqi Wang +1Multi-Hop QATool-Augmented Language Model Agents

  8. Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

    Jul 28, 2026Xinyi Hong, Pinjun Dong, Xinyang Yu +1Tool-Augmented Language Model AgentsLLM Tool Use

  9. E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

    Jul 26, 2026Weihuang Zheng, Tianyuan Zou, Eileen Ye +5Computer-Use Agent BenchmarksTool-Use Evaluation

  10. AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

    Jul 25, 2026Hao Jiang, Gangtao Xin, Yingdi Huang +35Tool-Augmented Language Model AgentsAgentic RL

  11. SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

    Jul 23, 2026Hao Zhang, Yiwen Zhao, Yixuan Zhang +2Tool-Augmented Language Model AgentsAudio Generation

  12. TerraLogic: A Benchmark for Hierarchical Geospatial Reasoning in Earth Observation

    Jul 14, 2026Yuhang Yan, Linchao Mou, Bokang Yang +1Spatial Reasoning BenchmarksHierarchical Planning

  13. ToFu: A White-Box, Token-Efficient Agent Harness for Researchers

    Jul 13, 2026Junhao Ruan, Yuan Ge, Bei Li +7Tool-Augmented Language Model AgentsSoftware Engineering Agents

  14. BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

    Jul 13, 2026Yuzhe Guo, Mengzhou Wu, Yuan Cao +4Tool-Augmented Language Model AgentsLanguage Model Generation Evaluation

  15. AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

    Jun 30, 2026Zhaojian Yu, Penghao Yin, Shuzheng Gao +3Tool-Augmented Language Model AgentsLLM Tool Use

  16. SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search

    Jun 30, 2026Ming Dai, Zhihong Lu, Jinjie Gu +5Tool-Augmented Language Model AgentsAgentic Search

  17. LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via State Proprioception

    Jun 29, 2026Binyan Xu, Haitao Li, Kehuan ZhangTool-Augmented Language Model AgentsLLM Agent Memory

  18. MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

    Jun 25, 2026Inderjeet Singh, Andrés Murillo, Motoyoshi Sekiya +2Multimodal RAGTool-Augmented Language Model Agents