LLM Agent Reliability

LLM: Large Language Model

Momentum

90 papers in the last four weeks, up 109% on the four weeks before. 0.9% of all new papers.

Jul 13Week of Sep 28

Latest papers 558

All topics
CardsList
  1. A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation

    Jul 24, 2026Fin Gentzen, Marla Grunewald, Iulisloi Zacarias +2Edge ComputingLanguage Model Calibration

  2. Reliability-Contagion Feasibility in LLM Multi-Agent Networks

    Jul 24, 2026Ruiwu Niu, Xincheng Shu, Ying ZhaoMulti-Agent LLM SystemsMulti-Agent Systems

  3. Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry

    Jul 23, 2026Natan Levy, Harel BergerAI Agent ReliabilityAI Agent Monitoring

  4. GuardianAgentBench: Where Agents Fail and How to Guard Them

    Jul 23, 2026Vishal Ishwar Naik, Chenyu Xu, Donna Dong +5AI Agent ReliabilityLLM Guardrails

  5. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

    Jul 23, 2026Pengyu Zhu, Lijun Li, Longju Yang +2Deep Research AgentsRAG Poisoning Attacks

  6. Same Game, Different Story: A Minimal Conservative Strategic Robustness Benchmark for Large Language Model Agents

    Jul 22, 2026Seyed Pouyan Mousavi Davoudi, Arshia Gharagozlou, Alireza Amiri-Margavi +2AI Agent BenchmarksFraming Effects in Language Models

  7. Agents in the Wild: Where Research Meets Deployment

    Jul 21, 2026Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz +3AI Agent ReliabilityMulti-Agent LLM Systems

  8. Operational Hallucination and Safety Drift in AI Agents

    Jul 20, 2026Shasha Yu, Fiona Carroll, Barry L. BentleyLLM Agent EvaluationAI Agent Safety

  9. Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking

    Jul 20, 2026Stefano Blando, Emanuele Guerrazzi, Riccardo Porcedda +3Agent-Based ModelingLLM Agents

  10. Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents

    Jul 20, 2026Yitao Wu, Si Shen, Rui Yang +2AI Agent ReliabilityLLM Self-Correction

  11. DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments

    Jul 19, 2026Jun Nie, Zhiqin Yang, Zhenheng Tang +4AI Agent ReliabilityDeep Research Agents

  12. Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution

    Jul 16, 2026Harris Borman, Herman Wandabwa, Fusun Yu +4Software Engineering AgentsBusiness Process Automation

  13. Tactile: Giving Computer-Using Agents Hands and Feet

    Jul 16, 2026Yong Liu, Zhenyi Zhong, Zhanpeng ShiComputer-Use AgentsGUI Agents

  14. DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

    Jul 15, 2026Huatao Li, Xinwei Geng, Yuheng Wang +9Computer-Use Agent BenchmarksLLM Agent Evaluation

  15. Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation

    Jul 15, 2026Dimple Vijay Kochar, Hae-Seung Lee, Anantha P. ChandrakasanElectronic Design AutomationLLM Agents

  16. Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents

    Jul 14, 2026Yaopei Zeng, Congchao Wang, JianHang Chen +3AI Agent ReliabilityLLM Agent Evaluation