LLM Agent Reliability

LLM: Large Language Model

Momentum

90 papers in the last four weeks, up 109% on the four weeks before. 0.9% of all new papers.

Jul 13Week of Sep 28

Latest papers 558

All topics
CardsList
  1. Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

    Oct 1, 2026Hao Wang, Ting HuangSmall Language ModelsAI Agent Benchmarks

  2. The Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents

    Oct 1, 2026David Fraile NavarroPersona ConsistencyPersonality Modeling in Language Models

  3. Revision-Aware Independent Agent Graphs for Dynamic Reasoning

    Oct 1, 2026Yan Luo, Selim-Antoine Lali, Jeremy Moebel +3Temporal ReasoningAI Agent Benchmarks

  4. Empty Commitments: When Agents Promise What They Cannot Deliver

    Oct 1, 2026Jiaqi Tang, Bingyu Shen, Lan Wei +6AI Agent ReliabilityLLM Agent Evaluation

  5. Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents

    Oct 1, 2026Junyu Guo, Shangding Gu, Ming Jin +1AI Agent AuditingCode Generation Evaluation

  6. PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

    Sep 30, 2026Yinghui He, Yapei Chang, Khushi Bhardwaj +4Language Model DistillationLLM Agent Training

  7. Learning When and How to Intervene: A Hindsight-Distilled Sentinel for Coding Agents

    Sep 30, 2026Jiangrui Zhao, Chenglong Li, Meng Zhang +1Coding AgentsRuntime Enforcement for AI Agents

  8. SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration

    Sep 30, 2026Weijie Ren, Yanwen Zhang, Hao Li +3Multi-Agent LLM SystemsQuestion Answering

  9. Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

    Sep 30, 2026Minghan Wang, Boyuan Wang, Jinhang Zuo +2LLM Agent ReliabilityLLM Agents

  10. EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning

    Sep 30, 2026Yitong Qiao, Yancheng Jin, Lei Liu +4Evidence-Grounded ReasoningAI Agent Benchmarks

  11. When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents

    Sep 30, 2026Shuyao Xiao, Shengling Wang, Xuan Chen +7Long-Horizon Agent EvaluationCausal Interventions in Language Models

  12. Coding Agents for Coding Theory

    Sep 30, 2026Abraham YeungCoding AgentsLLM Agent Reliability

  13. Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification

    Sep 30, 2026Yingfeng Luo, Shaowei Wei, Daixin Wang +7LLM Agent VerificationLLM Agent Reliability

  14. Multi-agent discussion gains less when dissent is withheld

    Sep 29, 2026Chand Sahil Mansuri, Xin Wang, Mengying Li +4Multi-Agent LLM SystemsOpinion Dynamics

  15. When Correct Memory Goes Wrong: Fuzzing Persistent Memory Use in LLM Agents

    Sep 29, 2026Yuqiao Meng, Luoxi Tang, Yingxue Zhang +2Agent MemoryLLM Agent Evaluation

  16. Absorbed in Inertia: Activation Analysis for Computer-Use Agents

    Sep 29, 2026Giulio Segalini, Zhi Wen Soi, Jérémie Decouchant +1Computer-Use AgentsLLM Agent Reliability

  17. AnyAct: Universal Action for Self-Evolving Agents

    Sep 29, 2026Lingrui Xu, Yangqin Jiang, Jiachang Zhang +2LLM Agent OrchestrationLLM Tool Use

  18. When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

    Sep 29, 2026Yaxin Gong, Gangyi Zhang, Chongming Gao +7Multi-Agent LLM SystemsMulti-Agent Collaboration

  19. Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models

    Sep 28, 2026Junru Zhu, Shiming Xie, Aime Lu Fan Chen +4Agent Failure AnalysisAI Agent Benchmarks

  20. BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation

    Sep 28, 2026Peilin Feng, Zhengyang Huang, Soujanya PoriaMulti-Agent CoordinationMemory-Augmented Agents

  21. MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

    Sep 28, 2026Xiaonan Xu, Wenjing WuLLM Tool UseModel Context Protocol

  22. Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses

    Sep 28, 2026Huachi Zhou, Yujing Zhang, Jiahe Du +7LLM Agent HarnessesLLM Agent Reliability

  23. Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias

    Sep 28, 2026Yinhong Liu, Zhili Tan, Zilin Wang +1LLM Agent EvaluationLLM Tool Use