Large Language Model Safety

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-21

27 new papers

A weekly snapshot of new work published in Large Language Model Safety.

Period ending 2026-09-14

12 new papers

A weekly snapshot of new work published in Large Language Model Safety.

Inside this field

Focused directions

855 papers

Latest in Large Language Model Safety

  1. Risky Business: Measuring The Faithfulness-Safety Tension

    Aug 4, 2026Dominik Meier, Luca Joshua Francis, Marco Bernhard Kaiser +3Large Language Model SafetyFaithfulness

  2. A Security-Oriented Lifecycle Model for Large Language Model Systems

    Aug 4, 2026Eleftherios Batzolis, George Drosatos, Vassilis Katsouros +1Large Language Model SafetySecurity

  3. LaCache: Robust Semantic Caching for LLM Serving

    Aug 3, 2026Jiacheng Liang, Yuhui Wang, Tanqiu Jiang +1CacheLanguage Model Evasion Attacks

  4. GPT-Red: Automated Red Teaming via Self-Play at Scale

    Jul 28, 2026Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15Red-TeamingLanguage Model Evasion Attacks

  5. Shieldstral

    Jul 28, 2026Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +273Content ModerationLarge Language Model Safety