Artificial Intelligence Safety

Recent momentum

-85%

6 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

4 new papers

A weekly snapshot of new work published in Artificial Intelligence Safety.

Period ending 2026-09-07

2 new papers

A weekly snapshot of new work published in Artificial Intelligence Safety.

231 papers

Latest in Artificial Intelligence Safety

Open your feed →
CardsList
  1. Rules or Character? Scaling Laws for AI Safety Design

    Aug 13, 2026Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu +1Artificial Intelligence SafetySafety

  2. Rethinking Agent Security as a Networking Problem

    Aug 12, 2026Van Tran, Taveesh Sharma, Tajveer Singh Dhesi +1SecurityArtificial Intelligence Safety

  3. Item Response Theory for AI Safety

    Aug 5, 2026Joshua Fonseca Rivera, Neil Shah, David Demitri Africa +1Item Response TheoryLarge Language Model Safety

  4. Evading Chain-of-Thought Monitoring Through Model Poisoning

    Aug 3, 2026Giorgio Severi, Shujaat Mirza, Blake Bullwinkel +1Reasoning TracesBackdoor Attacks

  5. A dataset of rated conceptual arguments

    Jul 29, 2026Emery Cooper, Caspar Oesterheld, Linh Chi Nguyen +2Abstract ArgumentationArtificial Intelligence Safety

  6. Agent Security Needs Redefinition through a Holistic Framework

    Jul 24, 2026Vincent Siu, Jingxuan He, Kyle Montgomery +3SecurityAuthorization

  7. Harmonizing AI Safety Thresholds

    Jul 17, 2026Wilber Sean Anterola, Matthew Ball, Luis F. Lafuerza +1Artificial Intelligence SafetyDecision Thresholds