cs.CRSep 7, 2026

An Empirical Measurement of Jailbreaking Evaluators

Authors: Yujie Mu

Organizations: Independent Researcher

Abstract

Expert evaluation of jailbreak responses is costly and difficult to scale, so the community increasingly relies on automated evaluators to determine whether an attack succeeds. However, jailbreak studies typically validate their chosen evaluator independently, repeatedly spending resources on similar evaluation efforts while making results across papers difficult to compare. Different evaluators also encode different definitions of jailbreak success, meaning that reported attack strength and apparent progress can depend substantially on which evaluator is used. We systematically compare six evaluators that recur in recent jailbreak attack and defense research: HarmBench, JailbreakBench, JailbreakRadar, StrongReject, JADES, and JailMeter. To our knowledge, no prior study has evaluated all six on the same human-labeled data under a controlled setup. We evaluate them on JailbreakQR and JailMeter-Eva, using human judgments as the reference, and measure agreement with humans, error types, and consistency across attack families. For evaluators that require a general-purpose LLM judge, we use a shared backbone to control for model-specific variation. We found that JADES exhibits the best overall performance, while HarmBench and StrongReject also demonstrate good performance.

Explore similar work

CardsList
  1. JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

    Jul 20, 2026Qingjia Huang, Jingyu Zhang, Jianguo Wu +6LLM Safety BenchmarksLLM Jailbreak Attacks

  2. GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods

    Feb 24, 2025Ruixuan Huang, Xunguang Wang, Zongjie Li +2Language Model Safety EvaluationLLM Jailbreak Attacks

  3. Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success

    May 9, 2026Carsten Maple, Abhishek Kumar, Riya TapwalLanguage Model Safety EvaluationJailbreak Attacks