Tool-Using Agents

Momentum

31 papers in the last four weeks, up 210% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 199

All topics
CardsList
  1. ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering

    Apr 20, 2026Yindong Zhang, Wenmian Yang, Yiquan Zhang +1LLM Agent OrchestrationAI Agent Benchmarks

  2. Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity

    Apr 19, 2026Leon Engländer, Sophia Althammer, Ahmet Üstün +2LLM Agent EvaluationTool-Using Agents

  3. Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories

    Apr 19, 2026Ivan Bercovich, Ivgeni Segal, Kexun Zhang +3Reward HackingAI Agent Security Benchmarks

  4. Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent Workflows

    Apr 17, 2026Luay Gharzeddine, Samer SaabMulti-Agent CoordinationMulti-Agent Orchestration

  5. VoxMind: An End-to-End Agentic Spoken Dialogue System

    Apr 17, 2026Tianle Liang, Yifu Chen, Shengpeng Ji +7Multi-Agent LLM SystemsSpoken Dialogue Systems

  6. El Agente Forjador: Task-Driven Agent Generation for Quantum Simulation

    Apr 16, 2026Zijian Zhang, Aiwei Yin, Amaan Baweja +4Multi-Agent LLM SystemsAI for Science

  7. Semantic Feature Analysis: Improving Agents Without Searching Over Rollouts

    Apr 12, 2026Yuval David, Fabiana Fournier, Lior Limonad +1LLM Agent Self-ImprovementTool-Using Agents

  8. Authoring for Living Worlds: Tool-Constrained LLM Agents for Executable Multi-Actor Scenarios

    Apr 11, 2026Nicolae Cudlenco, Mihai Masala, Marius LeordeanuLLM Agent OrchestrationTool-Using Agents

  9. Agentic Tool Use in Large Language Models

    Apr 1, 2026Jinchao Hu, Meizhi Zhong, Kehai Chen +2LLM Agent EvaluationTool-Using Agents

  10. Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

    Apr 1, 2026Alibek Kaliyev, Artem MaryanskyyAI Agent ReliabilityTool-Use Evaluation

  11. Formal Semantics for Agentic Tool Protocols: A Process Calculus Approach

    Mar 25, 2026Andreas SchlapbachAgent Communication ProtocolsFormal Verification

  12. LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

    Mar 20, 2026Xiang Long, Li Du, Yilong Xu +11Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  13. EvoGuard: An Extensible Agentic RL-based Framework for Practical and Evolving AI-Generated Image Detection

    Mar 18, 2026Chenyang Zhu, Maorong Wang, Jun Liu +2Ensemble LearningTool-Using Agents

  14. AgenTRIM: Tool Risk Mitigation for Agentic AI

    Jan 18, 2026Roy Betser, Amit Giloni, Shamik Bose +4LLM Agent SecurityIndirect Prompt Injection

  15. AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org

    Dec 12, 2025Jaehyung Lee, Justin Ely, Kent Zhang +3AI Agents for Scientific DiscoveryLLM Agent Evaluation

  16. PPTArena: A Benchmark for PowerPoint Editing

    Dec 2, 2025Michael Ofengenden, Yunze Man, Ziqi Pang +2LLM Agent EvaluationAI Agent Benchmarks

  17. Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

    Oct 18, 2025Vamshi Krishna Bonagiri, Ponnurangam Kumaragurum, Khanh Nguyen +1LLM Agent EvaluationTool-Using Agents

  18. El Agente Quntur: A research collaborator agent for quantum chemistry

    Date pendingJuan B. Pérez-Sánchez, Yunheng Zou, Jorge A. Campos-Gonzalez-Angulo +12AI Agents for Scientific DiscoveryQuantum Chemistry