AI Safety

Momentum

10 papers in the last four weeks, up 150% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 88

All topics
CardsList
  1. Muse Spark Safety & Preparedness Report

    May 14, 2026Cristina Menghini, Peter Ney, Hamza Kwisaba +117CybersecurityAI Safety Evaluation

  2. Mental Health AI Safety Claims Must Preserve Temporal Evidence

    May 9, 2026Srimonti Dutta, Ratna KandalaHealthcareMental Health

  3. Why Does Agentic Safety Fail to Generalize Across Tasks?

    May 7, 2026Yonatan Slutzky, Yotam Alexander, Tomer Slor +2AI Agent SafetyLLM Agent Safety

  4. Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems

    May 5, 2026Jie Zhou, Qin Chen, Liang HeAI AlignmentAI Safety

  5. Brainrot: Deskilling and Addiction are Overlooked AI Risks

    May 5, 2026Ilias Chalkidis, Anders SøgaardAI SafetyResponsible AI

  6. AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries

    May 2, 2026Wesley Shu, Peng WeiAI ControlAI Governance

  7. A Cellular Doctrine of Morality: Intrinsic Active Precision and the Mind-Reality Overload Dilemma

    May 2, 2026Ahsan AdeelAI Safety

  8. Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on Superintelligence

    May 2, 2026Edward Roussel, Lode Lauwaert, Torben Swoboda +4AI GovernanceAI Safety

  9. Open Problems in Frontier AI Risk Management

    Apr 28, 2026Marta Ziosi, Miro Plueckebaum, Stephen Casper +26AI Risk ManagementAI Governance

  10. Risk Reporting for Developers' Internal AI Model Use

    Apr 27, 2026Oscar Delaney, Sambhav Maheshwari, Joe O'Brien +2AI Risk ManagementAI Governance

  11. Agentic Microphysics: A Manifesto for Generative AI Safety

    Apr 16, 2026Federico Pierucci, Matteo Prandi, Marcantonio Bracale Syrnikov +2AI Agent SafetyAI Safety

  12. From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems

    Apr 2, 2026Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann +2AI SafetyFormal Verification

  13. TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

    Mar 31, 2026Qiucheng Yu, Ruijie Xu, Mingang Chen +2VLM EvaluationVLM Robustness

  14. Peer-Preservation in Frontier Models

    Mar 30, 2026Yujin Potter, Nicholas Crispino, Vincent Siu +2AI Agent SafetyLLM Safety Evaluation

  15. Defining Operational Conditions for Safety-Critical AI-Based Systems from Data

    Jan 29, 2026Johann Maximilian Christensen, Elena Hoemann, Frank Köster +1AI AssuranceAI Safety

  16. ECHO: A Participatory Framework for Bias-Anchored AI Harm Anticipation

    Nov 27, 2025Nicoleta Tantalaki, Sophia Vei, Athena VakaliAlgorithmic BiasHuman-Centered AI

  17. How Do Users Negotiate Harmful Value Conflicts with AI Companions? A Study with Minion, a Technology Probe for In-Situ Human-AI Conflict Response

    Nov 11, 2024Qing Xiao, Xianzhe Fan, Xuhui Zhou +4AI CompanionsHuman-AI Interaction

  18. From Silos to Systems: Process-Oriented Hazard Analysis for AI Systems

    Oct 29, 2024Shalaleh Rismani, Roel Dobbe, AJung MoonAI Risk ManagementAI Safety

  19. The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis

    Jul 22, 2024Benjamin Fresz, Vincent Philipp Göbels, Safa Omri +5Explainable Artificial IntelligenceAI Safety

  20. Why we need an AI-resilient society- Profiling Large Language Models

    Dec 18, 2019Thomas Bartz-Beielstein, Eva BartzAI SafetyAI Governance