Safety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profiles. Comparing models is more tractable at the level of individual attributes, yet it is often unclear whether even a single dataset's scores isolate any single attribute. One plausible candidate for such an attribute is harmful refusal, a model's tendency to refuse dangerous or policy-violating prompts. We examine whether it constitutes a single, measurable attribute in HELM Safety. Using a construct validity framework that stipulates that an attribute must exist before a test can measure it, we start with HELM Safety's four datasets that might plausibly target harmful refusal, but find that three are saturated. We subject the remaining dataset, HarmBench, to two psychometric tests to determine if a single attribute like harmful refusal could stand behind its score. First, multidimensional item response theory modeling strongly suggests that HarmBench does not measure a singular attribute. Second, a differential item functioning analysis finds items where models from different developers with the same refusal ability score differently. These flags largely disappear under scope-specific matching, a pattern consistent with aggregation effects but not sufficient to rule out domain-specific developer differences. Zooming out, HarmBench collapses distinct harm behaviors into one score, and the overall HELM safety aggregate further collapses HarmBench and scores from other datasets into a single top-line number. Any safety score that averages over datasets and items can hide saturation and conflate behaviors this way. We argue that a score should earn its single-attribute reading before models are compared with it.
Figures & tables
Figure 1: The HarmBench score supports a claim about a model only through a warrant. The warrant holds that one harmful refusal construct organizes the item responses. Our two construct validity tests probe that warrant, one from inside the response matrix and one from outside it.
Model
d
Params
Log-loss (SD)
Brier (SD)
AIC
BIC
Unidimensional (1D 2PL)
1
877
0.470 (0.003)
0.156 (0.001)
25,806
32,961
Unidimensional (1D 3PL)
1
1,275
0.322 (0.007)
0.096 (0.002)
17,600
28,001
Exploratory 2PL
2
1,754
0.287 (0.004)
0.087 (0.001)
16,112
30,421
Exploratory 2PL
3
2,631
0.284 (0.006)
0.086 (0.002)
17,782
39,245
Exploratory 2PL
4
3,508
0.285 (0.007)
0.086 (0.002)
19,490
48,108
Exploratory 2PL
5
4,385
0.282 (0.007)
0.085 (0.002)
21,329
57,100
Table 1: HarmBench model comparison across five repeated 80/20 response-level holdout splits. Log-loss and Brier entries are mean (SD) across splits after selecting one fit per model and split from twenty random starts by training log-likelihood; held-out responses were not used for restart selection. All models are 2PL except the unidimensional 3PL row included as the strongest unidimensional comparator. Exploratory models allow all items to load on all dimensions. Confirmatory 3D fixes standard, contextual, and copyright items to response-process factors. Confirmatory 7D fixes items to HarmBench harm-domain factors. Lower is better in the fit columns. Bold marks the best value in each column. AIC and BIC are means across splits and are rounded.
OpenAI vs. Anthropic
Closed vs. open
Scope
Items
MH
Logit
MH
Logit
Single score
398
13
17
0
0
3D response-process scopes
Standard
199
0
2
0
0
Contextual
99
1
0
0
0
Copyright
100
0
0
0
0
Table 2: Family-wise DIF flags after matching models on the relevant ability estimate. Rows include the single score, 3D response-process scopes, and 7D HarmBench-domain scopes. MH is the Mantel-Haenszel screen. Logit is the ridge-logistic sensitivity screen. Cutoffs are the 95th percentile of the maximum statistic across 5,000 bootstrap samples simulated with no group differences present.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Bernoulli-null eigenvalue check of the HarmBench item-correlation matrix. The orange line shows the random-data cutoff from Bernoulli matrices matched to HarmBench’s shape and item pass rates. Under one factor plus noise, only the first observed eigenvalue should exceed the cutoff. Several later eigenvalues do as well.
Model
Family
d
Params
Log-loss
Brier
AIC
BIC
Confirmatory 7D
2PL
7
1,363
0.255
0.077
13,961
25,080
Confirmatory 3D
2PL
3
1,039
0.258
0.078
13,707
22,183
Exploratory 3D
3PL
3
3,029
0.261
0.077
16,722
41,431
Confirmatory 3D
3PL
3
1,437
0.261
0.079
14,726
26,448
Exploratory 4D
3PL
4
3,906
0.262
0.076
17,873
49,737
Confirmatory 3D
1PL/Rasch
3
641
0.263
0.079
13,881
19,110
Appendix
Table 3: Model-family and content-partition robustness checks on the same five HarmBench train/test splits used in the main analysis. Params counts fitted item parameters and fitted model latent scores. Lower is better for held-out log-loss, Brier, AIC, and BIC.
Comparison
Group
Models
Mean score
Developer
OpenAI
21
0.866
Developer
Anthropic
11
0.889
Access type
Closed/API
48
0.776
Access type
Open-weight-like
33
0.507
Appendix
Table 4: Pre-specified DIF group comparisons. Means are raw HarmBench pass rates before ability matching.
Scope
Comparison
Items
MH item q95
MH FWER
Logit item q95
Logit FWER
Single score
OpenAI vs. Anthropic
398
94
13
102
17
Single score
Closed/API vs. open
398
41
0
51
0
Standard
OpenAI vs. Anthropic
199
27
0
25
2
Standard
Closed/API vs. open
199
13
0
12
0
Contextual
OpenAI vs. Anthropic
99
16
1
15
0
Contextual
Closed/API vs. open
99
8
0
8
0
Appendix
Table 5: DIF flags under Mantel-Haenszel and ridge-logistic parametric bootstraps. Item q95 counts use item-specific 95th-percentile null cutoffs. FWER q95 counts use the stricter max-statistic family-wise cutoff.
Domain
Items
Dev. MH
Dev. logit
Access MH
Access logit
Chemical/biological
56
0
0
0
0
Copyright
100
0
0
0
0
Cybercrime/intrusion
67
2
4
0
0
Harassment/bullying
25
0
0
0
0
General harmful content
22
0
0
0
0
Illegal behavior
63
0
0
0
0
Appendix
Table 6: Seven-domain DIF sensitivity. Entries are FWER q95 flag counts. The developer comparison is OpenAI versus Anthropic; the access-type comparison is closed/API versus open-weight-like.
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmarks. Second, psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, and roughly ten adaptively chosen items suffice for several individual benchmarks, cutting evaluation cost by 97-99%. Third, IRT supports audits of individual models, showing that it can be used to detect naive sandbagging and changes of model behind APIs. Overall, we show IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which we recommend frontier labs and evaluators adopt.
Joshua Fonseca Rivera, Neil Shah, David Demitri Africa +1
Independent · 1Independent · UK AI Security Institute +1
Safety benchmarks such as HarmBench rely on LLM judges to classify model responses as harmful or safe, yet the judge configuration, namely the combination of judge model and judge prompt, is typically treated as a fixed implementation detail. We show this assumption is problematic. Using a 2 x 2 x 3 factorial design, we construct 12 judge prompt variants along two axes, evaluation structure and instruction framing, and apply them using a single judge model, Claude Sonnet 4-6, producing 28,812 judgments over six target models and 400 HarmBench behaviors. We find that prompt wording alone, holding the judge model fixed, shifts measured harmful-response rates by up to 24.2 percentage points, with even within-condition surface rewording causing swings of up to 20.1 percentage points. Model safety rankings are moderately unstable, with mean Kendall tau = 0.89, and category-level sensitivity ranges from 39.6 percentage points for copyright to 0 percentage points for harassment. A supplementary multi-judge experiment using three judge models shows that judge-model choice adds further variance. Our results demonstrate that judge prompt wording is a substantial, previously under-examined source of measurement variance in safety benchmarking.
Xinran Zhang
University of California, Berkeley, Berkeley CA 94720, USA
Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both failure modes across 21 open-weight LLMs on four safety benchmarks (OR-Bench, XSTest, ToxiGen, BOLD), using a composition adjustment to isolate model sensitivity from dataset toxicity confounds. We report three findings. First, models adopt fundamentally different calibration strategies: conservative ecosystems such as Llama suppress unsafe outputs at the cost of elevated over-refusals, while permissive ecosystems such as DeepSeek and Qwen preserve helpfulness but tolerate higher harmful compliance. Second, demographic protection is unequal: models over-protect prominent racial and religious groups, frequently refusing even benign prompts about them, while providing substantially weaker protection against disability-targeted attacks. Third, refusal and compliance tendencies are stable within model families across generations and scales, suggesting that post-training objectives shape safety behavior more than architecture. Our results call for joint, demographically-aware, and multi-judge safety evaluation.
Alif Al Hasan, Sumon Biswas
Department of Computer and Data Sciences Case Western Reserve University Cleveland, OH, USA