LLM Agents

LLM: Large Language Model

Latest papers 433

All topics
CardsList
  1. What is Missing from AI Post-Training AI: An Empirical Analysis

    Aug 19, 2026Joy Jia Yin Lim, Xin Huang, Hao Peng +5AI Agent EvaluationLLM Agent Self-Improvement

  2. VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

    Aug 12, 2026Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder +6Tool-Use EvaluationAI Agent Benchmarks

  3. LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

    Aug 12, 2026Zhixin Zhang, Xinke Jiang, Zhibang Yang +5LLM-Guided RLLLM Agents

  4. Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration

    Aug 11, 2026Hunter McNichols, Kai Du, Andrew LanUser ModelingLLM Agents

  5. Mitigating Context Interference for Reliable and Efficient Search Agents

    Aug 11, 2026Boyang Xue, Bin Wu, Shuofei Qiao +8LLM AgentsLLM Context Management

  6. REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

    Aug 11, 2026Zixing Chen, Xingyuan Liu, Jie Zhu +6LLM Red TeamingLLM Agents

  7. RLMOpt: Adaptive Prompt Optimization via Recursive Language Models

    Aug 11, 2026Subhash Bangalore Satheesha, Nirvik Pande, Deepthi Duddempudi +1Large Language Model-Guided OptimizationLLM Agents

  8. SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

    Aug 10, 2026You Lu, Xinyu Huang, Bihuan Chen +1LLM Agent Skill LearningLLM Agents

  9. Agentic Router: An Execution-Grounded Continual Learning Approach With Memory

    Aug 10, 2026Yuxuan Chen, Rongpeng Li, Zhifeng Zhao +3Continual Learning for LLM AgentsLLM Agents

  10. Muscle Memory for Agents: Compile not Merely Retrieve

    Aug 10, 2026Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang +1Conversational MemoryLLM Agent Memory

  11. The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task

    Aug 9, 2026Marc Alier Forment, María José Casañ Guerrero, Francisco José García-Peñalvo +1AI Coding AgentsLLM Agent Evaluation

  12. Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

    Aug 8, 2026Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al AhsanConstrained DecodingLLM Agents

  13. SkillSmith: Enhancing Locally Deployed Agents via Automatic Skill Construction and Evolution

    Aug 8, 2026Xinle Jiang, Remy Xie, Ming TangLLM Agent Self-ImprovementLLM Agent Skill Learning

  14. When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment

    Aug 6, 2026Weihong Lin, Lin Sun, Xiangzheng ZhangLLM Agent EvaluationTool-Using Agents

  15. Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks

    Aug 5, 2026Nathan S Johnson, Ian AbshireAI Agents for Scientific DiscoveryLLM Agent Orchestration

  16. Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

    Aug 5, 2026Jinyi Han, Yuanjian Xu, Ying Liao +6LLM Agent EvaluationLLM Agent Skill Retrieval

  17. A game theory for foundation models shows new paths to rational cooperation through similarity inference

    Aug 4, 2026Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis +11Cooperative Game TheoryGame Theory

  18. Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure

    Aug 4, 2026Holly LewisLLM AgentsSelf-Evolving Agents

  19. Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory

    Aug 4, 2026Jakub Rada, Viliam LisýGame-Playing AgentsLLM Agents

  20. SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents

    Aug 3, 2026Yue Yao, Shengyuan Wang, Xin Chen +4Graph-Based RetrievalLLM Agent Skill Retrieval

  21. HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

    Aug 3, 2026Luan Zhang, Ruochen Zhou, Dandan Song +9Agent Harness OptimizationLLM Agents

  22. CRISP: Critical Step Perception for Training Efficient Deep Search Agents

    Aug 3, 2026Haosi Mo, Zihao Yan, Ruiqing Zhang +4LLM AgentsLLM Agent Training

  23. CT-PrepAgent: Bounded Policy and Controlled Execution for Adaptive CT Data Preparation

    Aug 2, 2026Xiaolin Fan, Yue Pei, Yingying Zhang +1LLM Agent OrchestrationLLM Agents

  24. Personalizing Large Language Model Agents with Small Policy Models

    Jul 31, 2026Dian Jin, Zhi Zhang, Huichao Li +3LLM PersonalizationLLM Agents