Large Language Model Safety

Recent momentum

-71%

27 papers in the last 28 days · 0.7% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

11 new papers

A weekly snapshot of new work published in Large Language Model Safety.

Period ending 2026-09-07

14 new papers

A weekly snapshot of new work published in Large Language Model Safety.

565 papers

Latest in Large Language Model Safety

Open your feed →
CardsList
  1. Inoculation Midtraining with Learned Neologisms

    Sep 14, 2026Kyle O'Brien, Edward James Young, Puria Radmard +4Large Language Model Safety

  2. Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

    Sep 14, 2026Mark Russinovich, Blake Bullwinkel, Giorgio Severi +2Large Language Model Safety

  3. Capability-Gated Language Models: Security Composes, Utility Does Not

    Aug 31, 2026Patrikas Vanagas, Augustas Mačijauskas, Laurynas LopataLarge Language Model SafetyAccess

  4. Item Response Theory for AI Safety

    Aug 5, 2026Joshua Fonseca Rivera, Neil Shah, David Demitri Africa +1Item Response TheoryLarge Language Model Safety