LLM Tool Use

LLM: Large Language Model

Momentum

30 papers in the last four weeks, up 233% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 185

All topics
CardsList
  1. Is Code Better Than Language for Algorithmic Reasoning

    Jun 14, 2026Terry Tong, Yu Feng, Surbhi Goel +1LLM EvaluationLLM Tool Use

  2. ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents

    Jun 13, 2026Rahul Suresh Babu, Laxmipriya Ganesh IyerAI Agent ReliabilityTool-Use Evaluation

  3. Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation

    Jun 13, 2026Kyle Gao, Joel Cumming, Jonathan Li +2Multi-Agent LLM SystemsLLM Tool Use

  4. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

    Jun 12, 2026Yongheng Zhang, Ziang Liu, Jiaxuan Zhu +17LLM Tool UseLong-Horizon LLM Agents

  5. HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents

    Jun 11, 2026Yaxin Du, Yifan Zhou, Yujie Ge +7Tool-Augmented Language Model AgentsLLM Agent Workflow Optimization

  6. Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents

    Jun 10, 2026Kushal Raj Bhandari, Ling Yue, Ching-Yun Ko +4Inference-Time SearchLLM Agent Workflow Optimization

  7. Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation

    Jun 9, 2026Yupu Hao, Zhuoran Jin, Huanxuan Liao +2Knowledge Augmentation for Language ModelsInference-Time Scaling

  8. QueryWeaver: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph Generation

    Jun 6, 2026Aishwarya Chakravarthy, Vidhi Kulkarni, Duen Horng ChauTool-Augmented Language Model AgentsTool-Use Planning

  9. MemToolAgent: Leveraging Memory for Tool Using Agents Based on Environment and User Feedback

    Jun 6, 2026Suleyman Armagan Er, Danilo Ribeiro, Yogesh Virkar +5Tool-Using AgentsLLM Tool Use

  10. Are Large Language Models Suitable for Graph Computation? Progress and Prospects

    Jun 5, 2026Yuting Zhang, Yi Han, Kai Wang +3LLM Reasoning with GraphsLLM Tool Use

  11. Translate-R1: Cost-Aware Translation Tool Use via Reinforcement Learning

    Jun 5, 2026Pratik Jayarao, Chaitanya Dwivedi, Himanshu Gupta +5RL for Language ModelsLLM Tool Use

  12. NTILC: Neural Tool Invocation via Learned Compression

    Jun 4, 2026Andrew Krikorian, Yayuan Li, Jason J. CorsoLLM Inference AccelerationLLM Tool Use

  13. ToolChoiceConfusion: Causal Minimal Tool Filtering for Reliable LLM Agents

    Jun 4, 2026Rahul Suresh Babu, Laxmipriya Ganesh IyerAI Agent ReliabilityTool-Use Evaluation

  14. When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

    Jun 4, 2026Dongsheng Zhu, Xuchen Ma, Yucheng Shen +5AI Agent ReliabilityAI Agent Benchmarks

  15. PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models

    Jun 2, 2026Zetian Ouyang, Linlin Wang, Gerard de Melo +1Numerical Reasoning in Language ModelsMathematical Reasoning Benchmarks

  16. Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions

    Jun 2, 2026Xuan Yang, Hao Xu, Tingfeng Hui +4LLM EvaluationTool-Use Evaluation

  17. AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design

    Jun 1, 2026Sahil Rahman, Maxx Richard RahmanRL for Language ModelsDirect Preference Optimization

  18. MAVEN: Improving Generalization in Agentic Tool Calling

    May 29, 2026Omkar Ghugarkar, Vishvesh Bhat, Muhammad Ahmed Mohsin +1Logical ReasoningLLM Agent Evaluation

  19. Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

    May 28, 2026Lorenz Kutschka, Bernhard GeigerTool-Augmented Language Model AgentsAI Agent Evaluation

  20. ParaTool: Shifting Tool Representations from Context to Parameters

    May 28, 2026Zekai Yu, Qi Meng, Qizhi Chu +3LLM Inference EfficiencyLLM Tool Use

  21. DeltaMCP: Incremental Regeneration via Spec-Aware Transformation for MCP servers

    May 27, 2026Aditya Pujara, Xiaogang Zhu, Hsiang-Ting ChenCode GenerationLLM Tool Use

  22. Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams

    May 25, 2026Tianda Sun, Dimitar KazakovLLM Tool UseLanguage Model Probing

  23. Memory-Induced Tool-Drift in LLM Agents

    May 24, 2026Mahavir Dabas, Jihyun Jeong, Ming Jin +1AI Agent BenchmarksLLM Tool Use

  24. HarnessAPI: A Skill-First Framework for Unified Streaming APIs and MCP Tools

    May 21, 2026Edwin JoseLLM Tool Use