Large Language Model Safety

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-21

27 new papers

A weekly snapshot of new work published in Large Language Model Safety.

Period ending 2026-09-14

12 new papers

A weekly snapshot of new work published in Large Language Model Safety.

Inside this field

Focused directions

855 papers

Latest in Large Language Model Safety

  1. Prefill Awareness in Large Language Models

    Jun 10, 2026Andy Wang, Parv Mahajan, David Demitri Africa +3Large Language Model SafetyFrontier Models

  2. Schützen: Evaluating LLM Safety in Bulgarian and German Contexts

    Jun 9, 2026Kiril Georgiev, Yuxia Wang, Dimitar Iliyanov Dimitrov +2Large Language Model SafetyGerman

  3. Do LLMs Make Neural Distinguishers Wise?

    Jun 9, 2026Tatsuya Sakagami, Masashi Hisai, Naoto YanaiLanguage Model Evasion AttacksText2Cypher

  4. Building Comparative Motivation Profiles with Instrumental Interventions

    Jun 6, 2026David Vella Zarb, Rustem Turtayev, Taywon Min +2DeceptionSycophancy

  5. Steering Vectors are an Adversarial Attack Surface

    Jun 4, 2026Abzal Aidakhmetov, Donato Crisostomi, Tommaso Mencattini +3Language Model Evasion AttacksActivation Steering