Artificial Intelligence Safety

Recent momentum

-66%

10 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

4 new papers

A weekly snapshot of new work published in Artificial Intelligence Safety.

Period ending 2026-09-07

2 new papers

A weekly snapshot of new work published in Artificial Intelligence Safety.

237 papers

Latest in Artificial Intelligence Safety

  1. A New Framework for Cybersecurity Refusals in AI Agents

    May 31, 2026Eliot Krzysztof Jones, Mateusz Dziemian, Matt Fredrikson +1CybersecurityArtificial Intelligence Safety

  2. Provably Secure Agent Guardrail

    May 28, 2026Benlong Wu, Weiming Zhang, Kejiang Chen +2Streaming GuardrailsArtificial Intelligence Safety

  3. Agent Security is a Systems Problem

    May 18, 2026Mihai Christodorescu, Earlence Fernandes, Ashish Hooda +11SecurityAgentic Systems

  4. Tracing Persona Vectors Through LLM Pretraining

    May 13, 2026Viktor Moskvoretskii, Dominik Glandorf, Jorge Medina Moreira +2PersonalityLarge Language Model Safety

  5. Conformity Generates Collective Misalignment in AI Agents Societies

    May 11, 2026Giordano De Marzo, Alessandro Bellina, Claudio Castellano +2Artificial Intelligence AlignmentConformity

  6. AI Native Asset Intelligence

    May 9, 2026Gal Engelberg, Leon Goldberg, Konstantin Koutsyi +3Artificial Intelligence Safety