Large Language Model Safety

Recent momentum

-47%

39 papers in the last 28 days · 0.6% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

13 new papers

A weekly snapshot of new work published in Large Language Model Safety.

Period ending 2026-09-14

11 new papers

A weekly snapshot of new work published in Large Language Model Safety.

Period ending 2026-09-07

14 new papers

A weekly snapshot of new work published in Large Language Model Safety.

532 papers

Latest in Large Language Model Safety

  1. AI Safety Training Can be Clinically Harmful

    Apr 25, 2026Suhas BN, Andrew M. Sherrill, Rosa I. Arriaga +2Mental Health SupportLarge Language Model Safety

  2. Training a General Purpose Automated Red Teaming Model

    Apr 24, 2026Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1Red-TeamingLanguage Model Evasion Attacks

  3. Estimating Tail Risks in Language Model Output Distributions

    Apr 24, 2026Rico Angell, Raghav Singhal, Zachary Horvitz +4Large Language Model Safety

  4. Detoxification for LLM: From Dataset Itself

    Apr 21, 2026Wei Shao, Yihang Wang, Gaoyu Zhu +4DetoxificationToxicity

  5. Surgical Repair of Insecure Code Generation in LLMs

    Apr 17, 2026Gustavo Sandoval, Brendan Dolan-Gavitt, Siddharth GargVulnerable CodeVulnerability Detection & Repair