AI Agent Reliability

Momentum

34 papers in the last four weeks, up 240% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 215

All topics
CardsList
  1. LACUNA: Safe Agents as Recursive Program Holes

    May 27, 2026Yaoyu Zhao, Yichen Xu, Oliver Bračevac +3AI Agent ReliabilityLLM Agents

  2. FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

    May 26, 2026Haoxuan Jia, Yang Liu, Bin Chong +10AI Agent ReliabilityLLM Guardrails

  3. Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems

    May 25, 2026Jianing Zhu, Yeonju Ro, John Robertson +5AI Agent ReliabilityAI Agent Benchmarks

  4. The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible

    May 25, 2026Lauri Lovén, Nam Do, Hassan Mehmood +2AI Agent ReliabilityAI Agent Governance

  5. Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

    May 18, 2026Leyao Wang, Yanan He, Peng Chen +5AI Agent ReliabilityLLM-as-a-Judge

  6. Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

    May 18, 2026Rishi Jha, Harold Triedman, Arkaprabha Bhattacharya +1Agent Failure AnalysisAI Agent Reliability

  7. Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

    May 18, 2026S. Bensalem, Y. Dong, M. Franzle +6AI Agent ReliabilityLLM Guardrails

  8. Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

    May 17, 2026Lecheng Yan, Ruizhe Li, Xicheng Han +5AI Agent ReliabilityLLM Agent Evaluation

  9. Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

    May 17, 2026Jinhu Qi, Muzhi Li, Jiahong Liu +9AI Agent ReliabilityAI Agent Evaluation

  10. ContractBench: Can LLM Agents Preserve Observation Contracts?

    May 17, 2026Jicheng Wang, Yifeng He, Zili Wang +3AI Agent ReliabilityLLM Agent Evaluation

  11. Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management

    May 16, 2026Carol Xuan Long, David Simchi-Levi, Feng Zhu +3AI Agent ReliabilityMulti-Agent Systems

  12. R2V Agent: Teaching SLMs When to Ask for Help

    May 15, 2026Raghu Vamshi Hemadri, Humaira Firdowse Mohammed, Rishabh Maheshwary +5AI Agent ReliabilityTool-Augmented Language Model Agents

  13. PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures

    May 15, 2026Yunan Lu, Luigi Liu, Omar Yahia +2AI Agent ReliabilityAdversarial Prompt Generation

  14. Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

    May 14, 2026Xiaohua Wang, Kai Yu, XuXiao Liang +2LLM Inference EfficiencyAI Agent Reliability

  15. Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy

    May 13, 2026Adarsh Kumarappan, Ananya MujooAI Agent ReliabilityMulti-Agent LLM Systems

  16. Useful Memories Become Faulty When Continuously Updated by LLMs

    May 13, 2026Dylan Zhang, Yanshan Lin, Zhengkun Wu +4AI Agent ReliabilityPersistent Memory for Language Models

  17. No Action Without a NOD: A Heterogeneous Multi-Agent Architecture for Reliable Service Agents

    May 12, 2026Zixu Yang, Hang Zheng, Nan Jiang +5AI Agent ReliabilityMulti-Agent LLM Systems

  18. Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability

    May 11, 2026Harsh Raj, Niranjan Orkat, Suvrorup Mukherjee +3AI Agent ReliabilityAgent Reliability