AI Agent Safety

Momentum

22 papers in the last four weeks, up 144% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 207

All topics
CardsList
  1. What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

    Jun 1, 2026Victor Ojewale, Suresh VenkatasubramanianAgent EvaluationAI Agent Safety

  2. Toward a Modular Architecture for Embedded AI Agent Systems at the Edge

    Jun 1, 2026Marcus Rüb, Michael GerhardsAI Agent SafetyAI Agent Governance

  3. SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence

    Jun 1, 2026Yuyan Bu, Haowei Li, Qirui Zheng +7LLM Agent EvaluationAI Agent Safety

  4. BraveGuard: From Open-World Threats to Safer Computer-Use Agents

    May 31, 2026Yunhao Feng, Xiaohu Du, Xinhao Deng +13LLM GuardrailsAI Agent Safety

  5. TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety

    May 30, 2026Zhepei Hong, Lin Wang, Liting Li +5AI Agent SafetyLLM Safety Evaluation

  6. ROGUE: Evaluating Corrigibility Failures in Frontier Computer-Use Agents

    May 29, 2026Jeremy Tien, Abishek Anand, Yu-Rou Tuan +3Agent ReliabilityComputer-Use Agents

  7. EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents

    May 29, 2026Dongwook Choi, Taeyoon Kwon, Bogyung Jeong +6LLM GuardrailsAI Agent Safety

  8. Gram: Assessing sabotage propensities via automated alignment auditing

    May 28, 2026David Lindner, Victoria Krakovna, Sebastian FarquharAI Coding AgentsAI Agent Auditing

  9. AI Loss of Control Incident Management: Response & Resilience

    May 28, 2026Ross GruetzemacherAI Risk ManagementAI Agent Safety

  10. Realistic honeypot evaluations for scheming propensity

    May 28, 2026Victoria Krakovna, David Lindner, Lewis Ho +2LLM Agent EvaluationAI Agent Safety

  11. Training Deliberative Monitors for Black-Box Scheming Detection

    May 28, 2026Aditya Sinha, Akshat Naik, Victor Gillioz +5AI Agent SafetyAI Agent Monitoring

  12. The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane

    May 27, 2026Tyler Akidau, Tyler Rockwood, Johannes Brüderl +1AI Agent SecurityAI Agent Safety

  13. Calibrating Conservatism for Scalable Oversight

    May 27, 2026William Overman, Mohsen BayatiScalable OversightAI Agent Safety

  14. Reasoning and Planning with Dynamically Changing Norms

    May 26, 2026Taylor Olson, Roberto Salas-Damian, Kenneth D. ForbusSymbolic PlanningAI Agent Safety

  15. Position: AI Safety Requires Effective Controllability

    May 26, 2026Yige Li, Yunhao Feng, Jun SunLLM Safety AlignmentAI Agent Safety

  16. Foundations of a Time-Consistent Counterfactual Actuarial Runtime for Autonomous AI Agents

    May 26, 2026Hao-Hsuan ChenCounterfactual EvaluationAI Risk Management

  17. Retrying vs Resampling in AI Control

    May 25, 2026James Lucassen, Adam KaufmanAI Agent SafetyAI Agent Monitoring

  18. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    May 21, 2026Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11LLM Safety BenchmarksComputer-Use Agent Benchmarks

  19. LLM Agents Make Collective Belief Dynamics Programmable: Challenges and Research Directions

    May 19, 2026Xin He, Junxi Shen, Yuchen Mou +4Opinion DynamicsAI Agent Safety

  20. Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

    May 18, 2026Rishi Jha, Harold Triedman, Arkaprabha Bhattacharya +1Agent Failure AnalysisAI Agent Reliability

  21. LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection

    May 18, 2026Lei Zhao, Abhay Bhaskar, Edgar DobribanAI Agent SecurityAI Agent Safety

  22. Ethical Hyper-Velocity (EHV): A Hardware-Rooted Zero-Trust Runtime Enforcement Architecture for Agentic AI Systems

    May 18, 2026Riddhi Mohan SharmaAI Agent SafetyRuntime Enforcement for AI Agents

  23. Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

    May 17, 2026Jinhu Qi, Muzhi Li, Jiahong Liu +9AI Agent ReliabilityAI Agent Evaluation

  24. TClone: Low-Latency Forking of Live GUI Environments for Computer-Use Agents

    May 17, 2026Yutong Huang, Vikranth Srivatsa, Alex Asch +2Computer-Use AgentsAI Agent Safety

  25. LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

    May 14, 2026Minbeom Kim, Lesly Miculicich, Bhavana Dalvi Mishra +6Continual Learning for LLM AgentsLLM Guardrails

  26. Conformity Generates Collective Misalignment in AI Agents Societies

    May 11, 2026Giordano De Marzo, Alessandro Bellina, Claudio Castellano +2Opinion DynamicsAI Agent Safety