LLM Agent Reliability

LLM: Large Language Model

Momentum

90 papers in the last four weeks, up 109% on the four weeks before. 0.9% of all new papers.

Jul 13Week of Sep 28

Latest papers 558

All topics
CardsList
  1. Diagnosing CFG Interpretation in LLMs

    Apr 22, 2026Hanqi Li, Lu Chen, Kai YuLLM EvaluationIn-Context Learning

  2. SWE-chat: Coding Agent Interactions From Real Users in the Wild

    Apr 22, 2026Joachim Baumann, Vishakh Padmakumar, Xiang Li +3Coding AgentsHuman-AI Collaboration

  3. Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

    Apr 21, 2026Bobo Li, Rui Wu, Zibo Ji +5AI Agent BenchmarksMulti-Agent Reasoning

  4. AI scientists produce results without reasoning scientifically

    Apr 20, 2026Martiño Ríos-García, Nawaf Alampara, Chandan Gupta +5AI Agent ReliabilityScientific Reasoning

  5. Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots

    Apr 20, 2026Shiquan Zhang, Tianyi Zhang, Le Fang +3Mobile GUI AutomationLLM Agent Evaluation

  6. Towards Self-Improving Error Diagnosis in Multi-Agent Systems

    Apr 19, 2026Jiazheng Li, Emine Yilmaz, Bei Chen +1Multi-Agent LLM SystemsMulti-Agent Systems

  7. Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity

    Apr 19, 2026Leon Engländer, Sophia Althammer, Ahmet Üstün +2LLM Agent EvaluationTool-Using Agents

  8. Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs

    Apr 19, 2026Carissa Cullen, Harry Garland, Alexander Roman +3Agent ReliabilityAI Agent Safety

  9. Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents

    Apr 19, 2026Hanlin Wang, Chak Tou Leong, Jian Wang +1Embodied AILLM Agent Reliability

  10. HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads

    Apr 18, 2026Justice Owusu Agyemang, Jerry John Kponyo, Obed Kwasi Somuah +3LLM Inference EfficiencyAI Coding Agents

  11. When Agents Go Quiet: Output Generation Capacity and Format-Cost Separation for LLM Document Synthesis

    Apr 17, 2026Justice Owusu Agyemang, Michael Agyare, Miriam Kobbinah +2LLM Inference EfficiencyLong-Form Document Generation

  12. Weak-Link Optimization for Multi-Agent Reasoning and Collaboration

    Apr 17, 2026Haoyu Bian, Chaoning Zhang, Jiaquan Zhang +4Multi-Agent CollaborationInference-Time Compute Allocation

  13. LLMs Corrupt Your Documents When You Delegate

    Apr 17, 2026Philippe Laban, Tobias Schnabel, Jennifer NevilleLLM ReliabilityLong-Horizon LLM Agents

  14. MOOSEnger: A Simulation-Aware AI Agent Framework for the MOOSE Ecosystem

    Mar 5, 2026Mengnan Li, Jason Miller, Zaid Abulawi +7AI Agent ReliabilityAgentic Code Generation

  15. A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development

    Mar 4, 2026Boyuan Guan, Wencong Cui, Levente JuhaszSoftware Engineering AgentsAI Coding Agents

  16. MASRubric: Auditing Information Flow in Multi-Agent Systems with Failure-Distilled Pitfall Rubrics

    Feb 26, 2026Yutong Wang, Siyuan Xiong, Xuebo Liu +4Agent Failure AnalysisMulti-Agent Systems

  17. When Is Enough Not Enough? Illusory Completion in Search Agents

    Feb 7, 2026Dayoon Ko, Jihyuk Kim, Sohyeon Kim +5AI Agent ReliabilityLLM Answer Verification