Safety

Recent momentum

-37%

25 papers in the last 28 days · 0.4% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

12 new papers

A weekly snapshot of new work published in Safety.

Period ending 2026-09-14

6 new papers

A weekly snapshot of new work published in Safety.

Period ending 2026-09-07

7 new papers

A weekly snapshot of new work published in Safety.

291 papers

Latest in Safety

  1. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    May 21, 2026Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11Large Language Model SafetySafety Evaluation

  2. Reinterpreting Safety Thresholds as Neuron Spiking Thresholds

    May 18, 2026Enrico Del Re, Mohamed Sabry, Cristina Olaverri-MonrealSpiking Neural NetworksAutonomous Driving

  3. Selective Safety Steering via Value-Filtered Decoding

    May 14, 2026Bat-Sheva Einbinder, Hen Davidov, Yee Whye Teh +2Large Language Model SafetySafety

  4. Auditing Agent Harness Safety

    May 14, 2026Chengzhi Liu, Yichen Guo, Yepeng Liu +8Agent HarnessRuntime Enforcement

  5. Shields to Guarantee Probabilistic Safety in MDPs

    May 11, 2026Linus Heck, Filip Macák, Roman Andriushchenko +2ShieldingSafety Constraints

  6. Conformity Generates Collective Misalignment in AI Agents Societies

    May 11, 2026Giordano De Marzo, Alessandro Bellina, Claudio Castellano +2Artificial Intelligence AlignmentConformity

  7. Why Does Agentic Safety Fail to Generalize Across Tasks?

    May 7, 2026Yonatan Slutzky, Yotam Alexander, Tomer Slor +2Agentic ControlSafety Constraints

  8. Safety Certification is Classification

    May 7, 2026Oliver Schön, Licio Romao, Sadegh SoudjaniCertificationDynamics Models