LLM Tool Use

LLM: Large Language Model

Momentum

30 papers in the last four weeks, up 233% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 185

All topics
CardsList
  1. To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling

    May 1, 2026Qinyuan Wu, Soumi Das, Mahsa Amani +5Tool-Use EvaluationLLM Tool Use

  2. Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents

    Apr 30, 2026Kaituo Zhang, Zhen Xiong, Mingyu Zhong +4Tool-Augmented Language Model AgentsAgentic Reasoning

  3. FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments

    Apr 28, 2026Amir Saeidi, Venkatesh Mishra, Souradeep Mukhopadhyay +4AI Agent ReliabilityMulti-Agent LLM Systems

  4. Uncertainty Quantification for LLM Function-Calling

    Apr 24, 2026Zihuiwen Ye, Lukas Aichberger, Michael Kirchhof +5Uncertainty QuantificationLLM Tool Use

  5. R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling

    Apr 22, 2026Aijia Cheng, Kailong Wang, Ling Shi +1RL for Language ModelsLLM Interpretability

  6. ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering

    Apr 20, 2026Yindong Zhang, Wenmian Yang, Yiquan Zhang +1LLM Agent OrchestrationAI Agent Benchmarks

  7. ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design

    Apr 18, 2026Yutang Ge, Guojiang Zhao, Sihang Li +7LLM Self-RefinementLLM Tool Use

  8. Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis

    Apr 17, 2026Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan +3Medical Image AnalysisLongitudinal Medical Image Analysis

  9. Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use

    Apr 17, 2026Ramit Pahwa, Apoorva Beedu, Parivesh Priye +4Audio-Language Model EvaluationVoice Agent Evaluation

  10. UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

    Apr 13, 2026Yijuan Liang, Xinghao Chen, Yifan Ge +5LLM Agent EvaluationAI Agent Benchmarks

  11. Agentic Tool Use in Large Language Models

    Apr 1, 2026Jinchao Hu, Meizhi Zhong, Kehai Chen +2LLM Agent EvaluationTool-Using Agents

  12. ZEBRAARENA: A Diagnostic Simulation Environment for Studying Reasoning-Action Coupling in Tool-Augmented LLMs

    Mar 19, 2026Wanjia Zhao, Ludwig Schmidt, Yejin Choi +3LLM EvaluationAI Agent Benchmarks

  13. CCTU: A Benchmark for Tool Use under Complex Constraints

    Mar 16, 2026Junjie Ye, Guoqiang Zhang, Wenjie Fu +4Tool-Use EvaluationAI Agent Benchmarks

  14. Tool Use Reduces Depth-Induced Collapse in OOD Reasoning

    Feb 24, 2026David Koplow, Tomer Galanti, Tomaso PoggioOOD GeneralizationLLM Tool Use

  15. LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning

    Jan 20, 2026Ye Tian, Zihao Wang, Onat Gungor +2HealthcareLong-Horizon Agent Evaluation

  16. Measuring Iterative Temporal Reasoning with Time Puzzles

    Jan 12, 2026Zhengxiang Wang, Zeyu DongLLM EvaluationTemporal Reasoning in Language Models

  17. Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models

    Jan 8, 2026Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani +3Multilingual Language Model EvaluationLLM Tool Use

  18. RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning

    Dec 31, 2025Xiang Gao, Yuguang Yao, Qi Zhang +5Knowledge Augmentation for Language ModelsLLM Tool Use

  19. AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning

    Dec 18, 2025Tzu-Han Lin, Wei-Lin Chen, Chen-An Li +3Retrieval-Augmented GenerationKnowledge Augmentation for Language Models

  20. Training Language Models to Use Prolog as a Tool

    Dec 8, 2025Niklas Mellgren, Peter Schneider-Kamp, Lukas Galke PoechReward HackingLogical Reasoning

  21. Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning

    Sep 27, 2025Ningning Xu, Yuxuan Jiang, Shubhashis Roy Dipta +1LLM Tool UseAlgorithmic Reasoning