LLM Agent Training

LLM: Large Language Model

Latest papers 359

All topics
CardsList
  1. Character Training for Risk-Averse Agents

    Sep 29, 2026Arav Dhoot, Punya Syon Pandey, Jamie Johnson +3Risk-Sensitive RLAI Agent Safety

  2. HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

    Sep 29, 2026Tongbo Chen, Junbo Niu, Zhengxi Lu +8Computer-Use AgentsGUI Agents

  3. AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

    Sep 29, 2026Hongjin Qian, Chaofan Li, Kun Luo +11LLM Self-RefinementLLM Agent Training

  4. AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction

    Sep 29, 2026Wentao Liu, Xi Chen, Siyu Song +13LLM Agent EvaluationLLM Agent Training

  5. SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

    Sep 29, 2026Renxi Wang, Mingshan Hee, Fajri Koto +2Supervised Fine-TuningLLM Agent Skill Learning

  6. Towards Communication-Efficient Social Intelligence in Language Agents

    Sep 28, 2026Linxiao Gong, Yijie Xu, Tianfu Wang +7LLM AgentsLLM Agent Training

  7. EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning

    Sep 28, 2026Shihan Dou, Shaofan Liu, Zhonghang Lu +8LLM Agent TrainingSelf-Improving Agents

  8. Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias

    Sep 28, 2026Yinhong Liu, Zhili Tan, Zilin Wang +1LLM Agent EvaluationLLM Tool Use

  9. Just-In-Time Agent Memory with Runtime Agentic Research

    Sep 28, 2026Bingyu Yan, Chaofan Li, Hongjin Qian +3Agent MemoryAgentic RAG

  10. When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

    Sep 27, 2026Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur +2Long-Horizon Agent EvaluationAI Agent Benchmarks

  11. ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments

    Sep 27, 2026Shengbin Yue, Hongru Wang, Siyuan Wang +3LLM Agent TrainingTool Use

  12. Breaking the Environment Wall: A Unified Framework for Preparing and Evolving Agent-Native Environments

    Sep 24, 2026Yukai Wu, Yuanjing Yang, Le Zhou +8LLM Agent Self-ImprovementLLM Agent Training

  13. From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

    Sep 24, 2026Xingyu Su, Abhishek Kumar, Qing Ping +5On-Policy Self-DistillationLLM Agent Training

  14. Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents

    Sep 24, 2026Liqin Ye, Haorui Wang, Fardin Ahmed +8LLM Agent EvaluationForecasting Benchmarks

  15. The Fellowship of the Query: Learning Retrieval Actions

    Sep 23, 2026Mohammed Al-Maamari, Saber Zerhoudi, Michael Granitzer +1LLM Fine-TuningAgentic RAG

  16. Agent-Editing World Model: Rethinking World Modeling for LLM Agents

    Sep 23, 2026Shuang Sun, Guoxin Chen, Fanzhe Meng +6LLM World ModelsWorld Models

  17. SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

    Sep 23, 2026Zhilong Ge, Yuting Shao, Yutao Yang +6Tool-Augmented Language Model AgentsLLM Agent Skill Learning

  18. Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

    Sep 23, 2026Xinjie Shen, Wei Fan, Xudong Guo +5Agentic RLReinforcement Learning

  19. Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

    Sep 22, 2026Laizhen Li, Jiarui Li, Juanjuan Zhao +4Agent Harness OptimizationLLM Agent Workflow Optimization

  20. The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

    Sep 22, 2026Wenbo Pan, Zhichao Liu, Shujie Liu +6Software Engineering BenchmarksLong-Horizon Agent Evaluation

  21. Harness-Zero: Harness Distillation via Agent-as-Harness

    Sep 21, 2026Haoran Ye, Yuxing Lu, Haonan Dong +2LLM Agent TrainingKnowledge Distillation

  22. Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning

    Sep 16, 2026Zhuo Chen, Zhen Zhang, Xinyu Wang +1LLM Inference EfficiencyLLM Agent Workflow Optimization

  23. SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

    Sep 15, 2026Md Tahmid Rahman Laskar, Xue-Yong Fu, Shashi Bhushan TNSupervised Fine-TuningLLM Tool Use