LLM Agent Reliability

LLM: Large Language Model

Momentum

90 papers in the last four weeks, up 109% on the four weeks before. 0.9% of all new papers.

Jul 13Week of Sep 28

Latest papers 558

All topics
CardsList
  1. Critic-Guided Heterogeneous Multi-Agent Reasoning for Reliable Mathematical Problem Solving

    Jun 4, 2026Muhammad Talha Sharif, Abdul RehmanLLM Self-CorrectionMulti-Agent Reasoning

  2. When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories

    Jun 3, 2026Avinash Baidya, Xinran Liang, Ruocheng Guo +2Weakly Supervised LearningLLM Agent Reliability

  3. A Taxonomy of Runtime Faults in Model Context Protocol Servers

    Jun 3, 2026Joshua Owotogbe, Indika Kumara, Willem-Jan van den Heuvel +3Agent Failure AnalysisModel Context Protocol

  4. Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

    Jun 3, 2026Haoyu Sun, Wenxuan Wang, Mingyang Song +5Tool-Augmented Language Model AgentsLLM Agent Evaluation

  5. AIP: A Graph Representation for Learning and Governing Agent Skills

    Jun 3, 2026Zachary Blumenfeld, Jim WebberAI Agent ReliabilityAgent Skill Learning

  6. SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language Models

    Jun 2, 2026Joel Sol, Homayoun NajjaranMulti-Agent CoordinationLLM Agent Evaluation

  7. Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

    Jun 2, 2026Ruida Wang, Jerry Huang, Pengcheng Wang +3Software Engineering AgentsFormal Verification

  8. Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents

    Jun 2, 2026Zhijie Ding, Weinan Hong, Zicheng Zhu +6Efficient Multimodal InferenceProactive Assistance

  9. SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence

    Jun 1, 2026Yuyan Bu, Haowei Li, Qirui Zheng +7LLM Agent EvaluationAI Agent Safety

  10. Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

    Jun 1, 2026Jiaming Wang, Ziteng Feng, Jiangtao Wu +8Agent Failure AnalysisDeep Research Agents

  11. QoEReasoner: An Agentic Reasoning Framework for Automated and Explainable QoE Diagnosis in RANs

    Jun 1, 2026Qizhe Li, Haolong Chen, Shan Dai +5LLM Agent OrchestrationWireless Communications

  12. Agent System Operations: Categorization, Challenges, and Future Directions

    Jun 1, 2026Zexin Wang, Changhua Pei, Yuanhao Liu +10Agent Failure AnalysisAI Agent Monitoring

  13. Self-Healing Agentic Orchestrators for Reliable Tool-Augmented Large Language Model Systems

    May 31, 2026Rahul Suresh Babu, Adarsh AgrawalLLM Agent OrchestrationLLM Agent Reliability

  14. Sophrosyne: Agentic Exploration of Relational Data Systems Needs Moderation

    May 29, 2026Madhav Jivrajani, Ramnatthan Alagappan, Aishwarya GanesanText-to-SQLTool-Using Language Agents

  15. Counterfactual Graph for Multi-Agent LLM Calibration

    May 28, 2026Jiatan Huang, Mingchen Li, Ziming Li +3AI Agent ReliabilityLanguage Model Calibration

  16. Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software

    May 28, 2026Nhat-Minh NguyenAI for ScienceAI Coding Agents

  17. Scaling Laws for Agent Harnesses via Effective Feedback Compute

    May 28, 2026Xuanliang Zhang, Dingzirui Wang, Keyan Xu +2Inference-Time ScalingAgent Harness Optimization

  18. GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

    May 28, 2026Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan +4LLM Agent Self-ImprovementLLM Agent Skill Learning

  19. Honest Lying: Understanding Memory Confabulation in Reflexive Agents

    May 28, 2026Prakhar Dixit, Sadia Kamal, Tim OatesAI Agent ReliabilityLLM Self-Refinement

  20. How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

    May 28, 2026Ningzhi Tang, Chaoran Chen, Gelei Xu +5Agent Failure AnalysisCoding Agents

  21. PatchBoard: Schema-Grounded State Mutation for Reliable and Auditable LLM Multi-Agent Collaboration

    May 28, 2026Shuyu Zhang, Yaqi Shi, Jiarui Zhang +2AI Agent ReliabilityMulti-Agent LLM Systems

  22. First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

    May 27, 2026Gianluca IngugliaAI Agents for Scientific DiscoveryAI Agent Evaluation