Tool-Using Language Agents

Momentum

9 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 44

All topics
CardsList
  1. MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation

    Jun 1, 2026Wenhao Wang, Peizhi Niu, Gongyi Zou +9LLM Agent EvaluationAI Agent Benchmarks

  2. Sophrosyne: Agentic Exploration of Relational Data Systems Needs Moderation

    May 29, 2026Madhav Jivrajani, Ramnatthan Alagappan, Aishwarya GanesanText-to-SQLTool-Using Language Agents

  3. SkillsInjector: Dynamic Skill Context Construction for LLM Agents

    May 28, 2026Yanchao Li, Wanhao Liu, Ben Gao +5Tool-Augmented Language Model AgentsLLM Agent Skill Learning

  4. Diagnosis Is Not Prescription: Linguistic Co-Adaptation Explains Patching Hazards in LLM Pipelines

    May 21, 2026Yoon Jeonghun, Kim DongchanLLM AgentsCausal Interventions in Language Models

  5. Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning

    May 18, 2026Yuval Shemla, Ayal Yakobe, Tanmay Agarwal +2Small Language ModelsLLM Tool Use

  6. RunAgent: Interpreting Natural-Language Plans with Constraint-Guided Execution

    May 1, 2026Arunabh Srivastava, Mohammad A., Khojastepour +2Multi-Agent LLM SystemsTool-Using Language Agents

  7. AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?

    May 1, 2026Ranit Karmakar, Jayita ChatterjeeComputer-Use Agent BenchmarksLLM Agent Evaluation

  8. ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship

    Apr 20, 2026Zhaopei Huang, Yanfeng Jia, Jiayi Zhao +3AI CompanionsAI Agent Benchmarks

  9. PersonalHomeBench: Evaluating Agents in Personalized Smart Homes

    Apr 18, 2026Manasa Bharadwaj, Yolanda Liu, InJung Yang +5AI Agent EvaluationMultimodal Agents

  10. Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis

    Apr 16, 2026Ke Zhang, Patricio Gallardo, Maziar Raissi +1Code TranslationLLM Agent Evaluation