LLM Tool Use

LLM: Large Language Model

Momentum

30 papers in the last four weeks, up 233% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 185

All topics
CardsList
  1. How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression

    Oct 7, 2026Xijie Gong, Tingxu Han, Jiahao Zhang +5LLM Tool UseLLM Decision-Making

  2. SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding

    Oct 5, 2026Zhi-Kai Chen, Song-Yan Li, De-Chuan Zhan +1Speculative DecodingLLM Inference Acceleration

  3. KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

    Oct 1, 2026Pengfei Li, Naufal Suryanto, Sicheng Zhang +1LLMs for CybersecurityLLM Tool Use

  4. The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching

    Oct 1, 2026Alessandro Pegoraro, Daryan Merx, Phillip Rieger +1LLM Agent SecurityPrivacy Leakage in Language Models

  5. Understanding Issues, Causes and Solutions in Open-Source LLM-based Multi-Agent Systems

    Oct 1, 2026Asad Ur Rehman, Syed Mohammad Kashif, Ruiyin Li +3Multi-Agent LLM SystemsMulti-Agent Orchestration

  6. AnyAct: Universal Action for Self-Evolving Agents

    Sep 29, 2026Lingrui Xu, Yangqin Jiang, Jiachang Zhang +2LLM Agent OrchestrationLLM Tool Use

  7. MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

    Sep 28, 2026Xiaonan Xu, Wenjing WuLLM Tool UseModel Context Protocol

  8. Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias

    Sep 28, 2026Yinhong Liu, Zhili Tan, Zilin Wang +1LLM Agent EvaluationLLM Tool Use

  9. PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents

    Sep 28, 2026Huayi Lai, Shichao Song, Qingchen Yu +4Long-Horizon Agent EvaluationAI Agent Benchmarks

  10. Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization

    Sep 28, 2026Xuanqi Zhang, Ruinan Jin, Running Yang +4Continual Learning for LLM AgentsLLM Agent Self-Improvement

  11. OpenFC: Learning Verification Policies towards Open-Search Fact Checking

    Sep 27, 2026Xinming Wang, Kaixiang Qiu, Yansong Lin +5RL for Language Model ReasoningAutomated Fact-Checking

  12. IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

    Sep 24, 2026Suvradip Paul, Chandra Bhushan, Harsh Sharma +4LLM Safety BenchmarksFinancial Services

  13. Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents

    Sep 23, 2026Jinqian Zhang, Haojun Xia, Shujiang Wu +4AI Agent SecurityLLM Agent Security

  14. Toolcompass: Guiding Tool Trialing, Not Suppressing It

    Sep 22, 2026Junlin Fang, Chong Zhang, Do Nguyen-Thanh +3Representation LearningTool-Augmented Language Model Agents

  15. SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

    Sep 15, 2026Md Tahmid Rahman Laskar, Xue-Yong Fu, Shashi Bhushan TNSupervised Fine-TuningLLM Tool Use

  16. What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity

    Sep 15, 2026Congjing Zhang, Vashishtha Patil, Henning Lange +1LLM GroundingLLM Pruning

  17. Towards a Deterministic Math Solver for Clinical Language Models

    Sep 12, 2026Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange +5Numerical Reasoning in Language ModelsLLM Reliability

  18. The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents

    Sep 10, 2026Bo Yan, Weikai Lin, Song WangTool-Using AgentsLLM Tool Use

  19. From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls

    Sep 10, 2026Hamed Jafarzadeh Asl, Yuanhao Yu, Vahid Partovi NiaLLM Inference EfficiencyOn-Device Language Model Inference