LLM Agent Reliability

LLM: Large Language Model

Momentum

90 papers in the last four weeks, up 109% on the four weeks before. 0.9% of all new papers.

Jul 13Week of Sep 28

Latest papers 558

All topics
CardsList
  1. Probing an Embodied LLM: When Higher Observation Fidelity Hurts Problem Solving

    May 19, 2026Oussama Zenkri, Oliver BrockLarge Language Model-Based Robot PlanningLLM Agent Evaluation

  2. Code as Agent Harness

    May 18, 2026Xuying Ning, Katherine Tieu, Dongqi Fu +39LLM Agent OrchestrationLong-Horizon LLM Agents

  3. AI for Auto-Research: Roadmap & User Guide

    May 18, 2026Lingdong Kong, Xian Sun, Wei Chow +17AI-Assisted Scientific ResearchScientific Workflow Automation

  4. STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery

    May 18, 2026Jiarui Su, Songjun Tu, Bei Sun +1LLM Self-RefinementAI Agents for Scientific Discovery

  5. VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

    May 17, 2026Hezhe Qiao, Hanghang Tong, Ee-Peng Lim +2Agent Failure AnalysisMulti-Agent LLM Systems

  6. Taming "Zombie'' Agents: A Markov State-Aware Framework for Resilient Multi-Agent Evolution

    May 17, 2026Taolin Zhang, Pukun Zhao, Qizhou Chen +5Multi-Agent LLM SystemsMulti-Agent Coordination

  7. ContractBench: Can LLM Agents Preserve Observation Contracts?

    May 17, 2026Jicheng Wang, Yifeng He, Zili Wang +3AI Agent ReliabilityLLM Agent Evaluation

  8. Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management

    May 16, 2026Carol Xuan Long, David Simchi-Levi, Feng Zhu +3AI Agent ReliabilityMulti-Agent Systems

  9. To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

    May 16, 2026Wei Shi, Ziheng Peng, Sihang Li +4Tool-Augmented Language Model AgentsCausal Interventions in Language Models

  10. R2V Agent: Teaching SLMs When to Ask for Help

    May 15, 2026Raghu Vamshi Hemadri, Humaira Firdowse Mohammed, Rishabh Maheshwary +5AI Agent ReliabilityTool-Augmented Language Model Agents

  11. PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures

    May 15, 2026Yunan Lu, Luigi Liu, Omar Yahia +2AI Agent ReliabilityAdversarial Prompt Generation

  12. The Scaling Laws of Skills in LLM Agent Systems

    May 15, 2026Charles Chen, Qiming Yu, Yuhang Gu +12Language Model Scaling LawsLLM Agent Reliability

  13. STAR: A Stage-attributed Triage and Repair framework for RCA Agents in Microservices

    May 15, 2026Junle Wang, Xingchuang Liao, Wenjun WuAgent Failure AnalysisRoot Cause Analysis

  14. Runtime-Structured Task Decomposition for Agentic Coding Systems

    May 14, 2026Shubhi Asthana, Bing Zhang, Chad DeLuca +2Software Engineering AgentsAgentic Workflows

  15. Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

    May 14, 2026Xiaohua Wang, Kai Yu, XuXiao Liang +2LLM Inference EfficiencyAI Agent Reliability

  16. Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use

    May 13, 2026Yize Cheng, Chenrui Fan, Mahdi JafariRaviz +2LLM Tool UseLLM Agent Reliability

  17. AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents

    May 13, 2026Hailin Zhong, Shengxin ZhuSoftware Engineering AgentsAgent Evaluation

  18. An Agentic LLM-Based Framework for Population-Scale Mental Health Screening

    May 13, 2026Giuliano Lorenzoni, Paulo Alencar, Donald CowanLLM Agent OrchestrationClinical NLP

  19. Useful Memories Become Faulty When Continuously Updated by LLMs

    May 13, 2026Dylan Zhang, Yanshan Lin, Zhengkun Wu +4AI Agent ReliabilityPersistent Memory for Language Models

  20. When Should an AI Workflow Release? Always-Valid Inference for Black-Box Generate-Verify Systems

    May 13, 2026Young Hyun Cho, Will Wei SunAnytime-Valid InferenceSequential Hypothesis Testing

  21. No Action Without a NOD: A Heterogeneous Multi-Agent Architecture for Reliable Service Agents

    May 12, 2026Zixu Yang, Hang Zheng, Nan Jiang +5AI Agent ReliabilityMulti-Agent LLM Systems

  22. CAX-Agent: A Lightweight Agent Harness for Reliable APDL Automation

    May 12, 2026Chenying Lin, Yichen Hai, Yi He +3LLM Agent OrchestrationLLM Agent Evaluation

  23. gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy

    May 11, 2026Tousif Islam, Digvijay Wadekar, Zihan ZhouGravitational-Wave AstronomyAI Coding Agents