Safety Benchmarks

Recent momentum

-75%

4 papers in the last 28 days · 0.1% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

2 new papers

A weekly snapshot of new work published in Safety Benchmarks.

Period ending 2026-09-07

3 new papers

A weekly snapshot of new work published in Safety Benchmarks.

91 papers

Latest in Safety Benchmarks

Open your feed →
CardsList
  1. Item Response Theory for AI Safety

    Aug 5, 2026Joshua Fonseca Rivera, Neil Shah, David Demitri Africa +1Item Response TheoryLarge Language Model Safety

  2. CRAX: Fast Safe Reinforcement Learning Benchmarking

    Jun 18, 2026Tristan Tomilin, Mourad Boustani, Mickey Beurskens +1Safe Reinforcement LearningSafety Benchmarks

  3. Configurable Reward Model for Balanced Safety Alignment

    May 28, 2026Zhengping Jiang, Mehran Khodabandeh, Akash Bharadwaj +5Safety AlignmentLarge Language Model Alignment

  4. Models That Know How Evaluations Are Designed Score Safer

    May 27, 2026Katharina Deckenbach, Haritz Puerto, Jonas Geiping +1Safety BenchmarksSafeguard

  5. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    May 21, 2026Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11Large Language Model SafetySafety Benchmarks