LLM Agent Reliability

LLM: Large Language Model

Momentum

90 papers in the last four weeks, up 109% on the four weeks before. 0.9% of all new papers.

Jul 13Week of Sep 28

Latest papers 558

All topics
CardsList
  1. How Strongly Should Task State Influence an LLM Agent?

    Sep 22, 2026Chenyu Zhang, Wonbin Kweon, Jiawei HanLLM Agent EvaluationRuntime Enforcement for AI Agents

  2. Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark

    Sep 22, 2026YanZe CaoLLM Agent EvaluationLLM Agent Reliability

  3. Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

    Sep 21, 2026Weihang Ding, Junfei ZhanLLM Agent EvaluationAI Agent Benchmarks

  4. ClashBench: Conflicts Leading Agents to Seize and Harm

    Sep 17, 2026Yuejin Xie, Yu Li, Dadi Guo +6AI Agent SecurityAI Agent Safety

  5. FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

    Sep 17, 2026Yanzhang Ma, Zhenghan Tai, Hanwei Wu +25LLM Agent Self-ImprovementLLM Agent Skill Learning

  6. Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

    Sep 16, 2026Mahsa Amani, Seungeon Lee, Abhisek Dash +9Tool-Augmented Language Model AgentsAgentic Search

  7. PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

    Sep 16, 2026Mika Okamoto, Ansel Kaplan ErolAI Agent AuditingLLM Agent Evaluation

  8. Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

    Sep 15, 2026Caiqi Zhang, Xiaochen Zhu, Chengzu Li +3Confidence Estimation in Language ModelsLLM Uncertainty Estimation

  9. Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

    Sep 14, 2026Yiwei Yang, Haoxiang Zhang, Bingbing Wen +6LLM Agent ReliabilityRL for Tool Use

  10. Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

    Sep 12, 2026Ruiqing Yue, Yu Cui, Zhuoyu Sun +11Agent Harness OptimizationLLM Agent Harnesses

  11. Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

    Sep 10, 2026Priyanka Mary Mammen, Emil Joswin, Srujananjali MedicherlaAI Agent MonitoringRepresentation Probing

  12. Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

    Sep 8, 2026Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta +2LLM Agent ReliabilitySelf-Evolving Agents

  13. The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution?

    Sep 8, 2026Boyang Wang, Yunhan Wang, Yalun WuLLM Agent EvaluationLLM Agent Reliability

  14. Vision: Data-Centric Anchoring for Robust and Interpretable Agentic AI

    Sep 8, 2026Arun Vignesh Malarkkan, Xinyuan Wang, Yanjie FuAI Agent ReliabilityDistribution Shift

  15. ResidualAuth: What Authorization State Must Language Agents Preserve under Revocable Delegation?

    Sep 8, 2026Moonwon Choi, Seokho Jeong, Sian Choi +1LLM Agent SecurityAccess Control

  16. Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning

    Sep 5, 2026Zhongan Bi, Qiwen Wang, Jianrong Jiang +11LLM Agent EvaluationWeb Search Agents

  17. FiMI Banking: A Sovereign Model for Indian Retail Banking

    Sep 3, 2026NPCI AI Research Team, Aman Kumar, Asit Desai +15Financial ServicesRL for Language Models

  18. KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

    Sep 3, 2026Yaxing Lyu, Shengjie Zhou, Binbin Toh +2Knowledge Conflicts in Language ModelsLLM Agent Evaluation

  19. Fresh Memory, Stale Plans: Derivation Currency for Distributed LLM-Agent Memory

    Sep 3, 2026Evan Chen, Shiqiang Wang, Christopher G. BrintonData ProvenanceRuntime Enforcement for AI Agents

  20. Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor

    Sep 2, 2026Rohith Reddy Bellibaltu, Manpreet Singh, Deepak Parashar +1Counterfactual EvaluationAlgorithmic Fairness

  21. MasterControl Seventeen Every Time

    Sep 2, 2026MasterControl AI LabLLM Agent EvaluationText-to-SQL