Artificial Intelligence Safety

Recent momentum

-66%

10 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

4 new papers

A weekly snapshot of new work published in Artificial Intelligence Safety.

Period ending 2026-09-07

2 new papers

A weekly snapshot of new work published in Artificial Intelligence Safety.

237 papers

Latest in Artificial Intelligence Safety

  1. Why Does Agentic Safety Fail to Generalize Across Tasks?

    May 7, 2026Yonatan Slutzky, Yotam Alexander, Tomer Slor +2Agentic ControlSafety Constraints

  2. Open Problems in Frontier AI Risk Management

    Apr 28, 2026Marta Ziosi, Miro Plueckebaum, Stephen Casper +26Artificial Intelligence SafetyData-Driven

  3. AI Safety Training Can be Clinically Harmful

    Apr 25, 2026Suhas BN, Andrew M. Sherrill, Rosa I. Arriaga +2Mental Health SupportLarge Language Model Safety

  4. Agentic Microphysics: A Manifesto for Generative AI Safety

    Apr 16, 2026Federico Pierucci, Matteo Prandi, Marcantonio Bracale Syrnikov +2Artificial Intelligence SafetyAgentic Systems

  5. Peer-Preservation in Frontier Models

    Mar 30, 2026Yujin Potter, Nicholas Crispino, Vincent Siu +2Frontier ModelsArtificial Intelligence Alignment

  6. LLMs Encode Harmfulness and Refusal Separately

    Jul 16, 2025Jiachen Zhao, Jing Huang, Zhengxuan Wu +2Large Language Model SafetyRefusals