Period ending 2026-09-21
7 new papers
A weekly snapshot of new work published in Higher Attack Success Rates.
Twelve weeks of publication activity for this topic as it is defined today.
Weekly history
What was published in this field, kept on the site without email delivery.
Period ending 2026-09-21
A weekly snapshot of new work published in Higher Attack Success Rates.
Period ending 2026-09-14
A weekly snapshot of new work published in Higher Attack Success Rates.
Period ending 2026-09-07
A weekly snapshot of new work published in Higher Attack Success Rates.
Inside this field
Within Higher Attack Success Rates
Within Higher Attack Success Rates
Within Higher Attack Success Rates
Within Higher Attack Success Rates
Within Higher Attack Success Rates
Within Higher Attack Success Rates
235 papers
Llama, Gemma, and Qwen model families, RAS separates aligned models from uncensored and abliterated variants, tracks output-level attack success rate, and is substantially faster than judge-based evaluation. These results suggest that refusal alignment provides a compact and efficient signal for white-box LLM safety evaluation.