Period ending 2026-09-21
16 new papers
A weekly snapshot of new work published in Safety Alignment.
Twelve weeks of publication activity for this topic as it is defined today.
Weekly history
What was published in this field, kept on the site without email delivery.
Period ending 2026-09-21
A weekly snapshot of new work published in Safety Alignment.
Period ending 2026-09-14
A weekly snapshot of new work published in Safety Alignment.
Period ending 2026-09-07
A weekly snapshot of new work published in Safety Alignment.
Inside this field
Within Safety Alignment
Within Safety Alignment
Within Safety Alignment
Within Safety Alignment
464 papers
Llama, Gemma, and Qwen model families, RAS separates aligned models from uncensored and abliterated variants, tracks output-level attack success rate, and is substantially faster than judge-based evaluation. These results suggest that refusal alignment provides a compact and efficient signal for white-box LLM safety evaluation.