LLM Agent Safety

LLM: Large Language Model

Momentum

24 papers in the last four weeks, up 50% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 193

All topics
CardsList
  1. A Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM Agents

    Jun 30, 2026Javal Vyas, Milapji Singh Gill, Artan Markaj +2Language Model-Based ControlFault-Tolerant Control

  2. A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control

    Jun 29, 2026Idelfonso B. R. Nogueira, Sigurd SkogestadMulti-Agent LLM SystemsLanguage Model-Based Control

  3. Entity Binding Failures in Tool-Augmented Agents

    Jun 29, 2026Rahul Suresh Babu, Shashank IndukuriTool-Using AgentsLLM Agent Safety

  4. Agent Safety Is Action Alignment

    Jun 27, 2026Shawn Li, Yue ZhaoLanguage Model Safety EvaluationAI Alignment

  5. It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents

    Jun 26, 2026Yiming Sun, Chen Chen, Zifan Zhou +1Computer-Use AgentsMobile GUI Agents

  6. GUI agent: Guided Exploration of User-Sensitive Screens

    Jun 24, 2026Aradhana Nayak, Mussadiq Nazeer, Wang Peng +1GUI AgentsHuman-in-the-Loop AI

  7. AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming

    Jun 23, 2026Pingchuan Ma, Zhaoyu Wang, Zimo Ji +5LLM GuardrailsProgram Synthesis

  8. Agent MechSuits: Mechanistic Subspace Safety Steering for Multi-Turn CLI Agents

    Jun 21, 2026Weidi Luo, Qiming Zhang, Yihao Quan +5Computer-Use AgentsLLM Interpretability

  9. PrivacyAlign: Contextual Privacy Alignment for LLM Agents

    Jun 19, 2026Manveer Singh Tamber, Abhay Puri, Marc-Etienne Brunet +3Reward ModelingPrivacy Leakage in Language Models

  10. Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention

    Jun 19, 2026Chubin Zhang, Zhenglin Wan, Xingrui Yu +5LLM Agent Safety

  11. LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

    Jun 18, 2026Md Nayem Uddin, Amir Saeidi, Eduardo Blanco +1Tool-Augmented Language Model AgentsCustomer Support Automation

  12. NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

    Jun 18, 2026Hanwool Lee, Dasol Choi, Bokyeong Kim +2Multi-Agent LLM SystemsSafety-Critical Control

  13. Phoenix: Safe GitHub Issue Resolution via Multi-Agent LLMs

    Jun 18, 2026Kipngeno Koech, Muhammad Adam, Baimam Boukar Jean Jacques +1Multi-Agent LLM SystemsSoftware Engineering Agents

  14. Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications

    Jun 16, 2026Divyansh Srivastava, Shreya Ghosh, Anshul Verma +1Multi-Agent LLM SystemsLLM Hallucination Mitigation

  15. Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation

    Jun 13, 2026Kyle Gao, Joel Cumming, Jonathan Li +2Multi-Agent LLM SystemsLLM Tool Use

  16. Resilient Consensus in Agentic AI

    Jun 12, 2026Sribalaji C. Anand, George J. PappasMulti-Agent LLM SystemsMulti-Agent Consensus

  17. Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH

    Jun 12, 2026Vikhyath Kothamasu, Virginia Smith, Chhavi YadavAI Agent SafetyAI Agent Safety Benchmarks

  18. Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval

    Jun 9, 2026Jiandong Ding, Honglei Ji, Ming Liu +1Agentic RetrievalLLM Agent Skill Retrieval

  19. ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

    Jun 7, 2026Lu Jia, Haibo Tong, Feifei Zhao +5AI Agent Safety BenchmarksLLM Agent Safety

  20. SciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery Agents

    Jun 6, 2026Tanush Swaminathan, Runmin Jiang, Letian Zhang +1AI Agents for Scientific DiscoveryLLM Safety Evaluation