CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks
Organizations: Independent researcher
Abstract
An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label, matching the trivial always-first baseline at 0.581. Judges with the same parameter count score far higher everywhere: a reward model reaches 0.764 to 0.976 and a generative judge 0.611 to 0.778, and every gap to CLM is significant after Benjamini-Hochberg correction. Two properties do work. Raw confidences are overconfident by up to +0.401, yet one pooled temperature fit on held-out calibration items repairs expected calibration error to at most 0.062, and the repaired confidence ranks the model's own errors above chance on three of six benchmarks. The decision order-flip rate is 0.0002 against 0.2188 for the generative judge, and the length-preference shift is -0.023 against -0.217. The confidence-gated cascade, however, escalates between 0.923 and 1.000 of items to the strong judge at the preregistered 0.97 retention bar: calibrated confidence about a near-chance judge has almost nothing to keep. The design: five public preference benchmarks and one hallucination benchmark with real labels, scored under a preregistration frozen before any test item was seen, against generative, reward-model, and trivial baselines, with per-item predictions released.
Figures & tables
| Benchmark | CLM (2048) | laya (512) | Generative (4096) | RM (4096) |
|---|---|---|---|---|
| RewardBench | 0.01 / 0.99 | 0.40 / 0.76 | 0.00 / 1.00 | 0.00 / 1.00 |
| RewardBench 2 | 0.56 / 0.83 | 0.93 / 0.27 | 0.09 / 1.00 | 0.06 / 1.00 |
| RewardBench 2 pairs | 0.06 / 1.00 | 0.77 / 0.50 | 0.00 / 1.00 | 0.00 / 1.00 |
| RM-Bench | 0.05 / 0.99 | 0.84 / 0.51 | 0.01 / 1.00 | 0.00 / 1.00 |
| JudgeBench | 0.15 / 0.97 | 0.99 / 0.42 | 0.06 / 0.99 | 0.06 / 0.99 |
| HaluEval | 0.19 / 0.71 | 0.25 / 0.54 | 0.19 / 0.71 | 0.19 / 0.71 |
| Judge | RewardBench | RewardBench 2 | RB2 pairs | RM-Bench | JudgeBench | HaluEval |
|---|---|---|---|---|---|---|
| CLM | 0.549 | 0.351 | 0.593 | 0.508 | 0.475 | 0.581 |
| CLM (tuned heads) | 0.521 | 0.387 | 0.647 | 0.539 | 0.535 | 0.581 |
| laya | 0.512 | 0.280 | 0.526 | 0.494 | 0.498 | 0.544 |
| Qwen3-8B (generative) | 0.778 | 0.675 | 0.758 | 0.676 | 0.611 | 0.728 |
| Qwen3-32B (strong, hosted) | 0.868 | 0.732 | 0.827 | 0.726 | 0.615 | 0.745 |
| Skywork-Reward-V2-8B | 0.976 | 0.876 | 0.950 | 0.923 | 0.764 | — |
| Benchmark | Identity ECE | Signed gap | Pooled | Per-bench | Isotonic |
|---|---|---|---|---|---|
| RewardBench | 0.235 | +0.235 | 0.032 | 0.032 | 0.015 |
| RewardBench 2 | 0.402 | +0.401 | 0.061 | 0.023 | 0.060 |
| RewardBench 2 pairs | 0.283 | +0.283 | 0.049 | 0.027 | 0.020 |
| RM-Bench | 0.343 | +0.343 | 0.042 | 0.042 | 0.028 |
| JudgeBench | 0.354 | +0.351 | 0.062 | 0.083 | 0.122 |
| HaluEval | 0.255 | +0.254 | 0.060 | 0.009 | 0.040 |
| Benchmark | CLM | Strong | Agree | Cascade | Escal. | Retained | Random |
|---|---|---|---|---|---|---|---|
| RewardBench | 0.549 | 0.868 | 0.532 | 0.854 | 0.933 | 0.983 | 0.848 |
| RewardBench 2 | 0.351 | 0.732 | 0.339 | 0.732 | 1.000 | 1.000 | 0.732 |
| RewardBench 2 pairs | 0.593 | 0.827 | 0.580 | 0.818 | 0.923 | 0.989 | 0.809 |
| RM-Bench | 0.508 | 0.760 | 0.475 | 0.741 | 0.954 | 0.976 | 0.749 |
| JudgeBench | 0.475 | 0.615 | 0.482 | 0.615 | 0.979 | 1.000 | 0.613 |
| HaluEval | 0.581 | 0.745 | 0.662 | 0.745 | 1.000 | 1.000 | 0.745 |
| Benchmark | laya | CLM | Skywork-Reward-V2-8B | Qwen3-8B (generative) |
|---|---|---|---|---|
| RewardBench | 23 | 338 | 244 | 746 |
| RewardBench 2 | 29 | 674 | 558 | 1,058 |
| RewardBench 2 pairs | 24 | 233 | 359 | 898 |
| RM-Bench | 22 | 106 | 230 | 780 |
| JudgeBench | 32 | 292 | 331 | 937 |
| HaluEval | 23 | 63 | — | 635 |