RL for Tool Use

RL: Reinforcement Learning

Momentum

17 papers in the last four weeks, up 42% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 127

All topics
CardsList
  1. SEAL: Synergistic Co-Evolution of Agents and Learning Environments

    May 23, 2026Yihao Hu, Zhihao Wen, Xiujin Liu +3LLM Agent TrainingSelf-Improving Agents

  2. B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation

    May 22, 2026Mario Markov, Stefan Maria Ailuro, Mohammad Mahdi +2Image SegmentationReinforcement Learning

  3. ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

    May 19, 2026Zuhao Yang, Kaichen Zhang, Sudong Wang +7Tool Use in VLMsAgentic RL

  4. Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning

    May 19, 2026Qinghe Ma, Zhen Zhao, Yiming Wu +3Multimodal Large Language ModelsMultimodal Reasoning

  5. EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL

    May 18, 2026Minrui Xu, Zilin Wang, Mengyi DENG +12Tool-Augmented Language Model AgentsAgentic RL

  6. Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

    May 17, 2026Yuxuan Lu, Ziyi Wang, Yingzhou Lu +12Tool-Augmented Language Model AgentsSynthetic Data Generation

  7. Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use

    May 14, 2026Renning Pang, Tian Lan, Leyuan Liu +3Language Model CalibrationLLM Tool Use

  8. Reinforcement Learning for Tool-Calling Agents in Fast Healthcare Interoperability Resources (FHIR)

    May 13, 2026Marius S. Knorr, Robert Müller, Jan P. Bremer +1HealthcareRL for Language Model Reasoning

  9. ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents

    May 12, 2026Xuhao Hu, Xi Zhang, Haiyang Xu +6Agentic RLRL for GUI Agents

  10. When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents

    May 12, 2026Xiaolin Zhou, Aojie Yuan, Zheng Luo +12Domain RandomizationAI Agent Benchmarks

  11. GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

    May 12, 2026Sijia Li, Yuchen Huang, Zifan Liu +7RL for Language Model ReasoningCredit Assignment in RL

  12. AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design

    May 9, 2026Haoze Lv, Ning Lu, Ziang Zhou +1Agentic RLAutomated Heuristic Design

  13. CoCoDA: Co-evolving Compositional DAG for Tool-Augmented Agents

    May 8, 2026Ziyang Yu, Qiyue Li, Liang ZhaoLLM Tool UseLLM Agent Training

  14. Learning CLI Agents with Structured Action Credit under Selective Observation

    May 8, 2026Haoyang Su, Ying WenComputer-Use AgentsCredit Assignment in RL

  15. Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning

    May 7, 2026Qianjia Cheng, Yuchen Zhang, Zhilin Wang +9RL for Language Model ReasoningRL for Tool Use

  16. A2^2TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping

    May 7, 2026Dingwei Chen, Zefang Zong, Zhipeng Ma +5Credit Assignment in RLPolicy Gradient Methods

  17. Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL

    May 6, 2026Yaxun Dai, Baolin Sun, Junying Wang +6Credit Assignment in RLText-to-SQL

  18. S3S^3-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data

    May 2, 2026Harsh Goel, Akhil Udathu, Susmija Jabbireddy +2Multi-Hop QASynthetic Data Generation

  19. Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model

    Apr 20, 2026Chenming Tang, Hsiu-Yuan Huang, Weijie Liu +3LLM Agent TrainingRL for Tool Use

  20. Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis

    Apr 16, 2026Zhiyuan Zhai, Wenjing Yan, Xiaodan Shao +1LLM Agent TrainingRL for LLM Agents

  21. Agentic Tool Use in Large Language Models

    Apr 1, 2026Jinchao Hu, Meizhi Zhong, Kehai Chen +2LLM Agent EvaluationTool-Using Agents

  22. PhGPO: Pheromone-Guided Policy Optimization for Long-Horizon Tool Planning

    Feb 14, 2026Yu Li, Guangfeng Cai, Shengtian Yang +5LLM-Guided RLTool-Use Planning

  23. Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

    Feb 1, 2026Yu Li, Mingyang Yi, Xiuyu Li +6RL for Language Model ReasoningGradient Interference

  24. PaperScout: An Autonomous Agent for Academic Paper Search with Process-Aware Sequence-Level Policy Optimization

    Jan 15, 2026Tingyue Pan, Jie Ouyang, Mingyue Cheng +7Agentic RLPolicy Optimization