Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
Organizations: Griffith University · Nanyang Technological University · UNSW · Independent Researcher · Deakin University · George Mason University · Wake Forest University
Abstract
Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44 benchmarks and five target models, labelled by each benchmark's scorer and, on two, by humans. Many of these failures are relational, defined against a reference, such as the user's belief or an injected instruction, that the response alone does not reveal. Our key idea is therefore to vary what Jev is asked separately from what it sees: the question's wording and answer type on one side, the fields of the input on the other. A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mostly through fields that encode the label. Jev matches the reference scorer's agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers. Code and data: https://github.com/sumleo/RLCDAlignBench.
Figures & tables
| Benchmarks | # | R/J/M | V/D/U | Items |
| Sycophancy (Qwen3.5-2B): the reply defers to the user’s stated view or flips its verdict to please the user. | ||||
| ELEPHANT ( Cheng et al., 2026 ) , SYCON-Bench ( Hong et al., 2025 ) , SycophancyEval 2 ( Sharma et al., 2024 ) | 4 | 1/2/1 | 0/2/2 | 639 |
| Jailbreaks (Phi-4-mini): the reply complies with a harmful (jailbroken) request. | ||||
| HarmBench ( Mazeika et al., 2024 ) , JailbreakBench 2 ( Chao et al., 2024 ; Chao et al., 2025 ) , StrongREJECT ( Souly et al., 2024 ) | 4 | 0/4/0 | 1/0/3 | 414 |
| Deception (Gemma-2-2B): a statement contradicts the model’s own elicited belief or serves a deceptive goal. | ||||
| DeceptionBench † ( Huang et al., 2025 ) , MASK 3 ( Ren et al., 2025 ) | 4 | 0/4/0 | 0/0/4 | 540 |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Split | Scorer | Label source | Label | Pos. | Rate | Status | |
|---|---|---|---|---|---|---|---|---|
| Sycophancy (target model: Qwen3.5-2B) | ||||||||
| ELEPHANT (AITA) ( Cheng et al., 2026 ) | H | Rule | rule (verdict tokens) | unval. | 255 | 134 | 0.53 | ok |
| SYCON-Bench (false premise) ( Hong et al., 2025 ) | HO | Multi-turn | GPT-4o, per turn | unval. | 17 | 14 | 0.82 | degen. (min. 3) |
| SycophancyEval (answer) ( Sharma et al., 2024 ) | H | Judge | GPT-4o | defect | 150 | 123 | 0.82 | ok |
| SycophancyEval (feedback) ( Sharma et al., 2024 ) | H | Judge | GPT-4o | defect | 217 | 30 | 0.14 | ok |
| Jailbreaks (target model: Phi-4-mini) | ||||||||
| Family | What Jev is asked | Score |
|---|---|---|
| Direct Noul (minimal, statement, definition, criteria) | The benchmark-specific yes/no judgment in four wordings: minimal , statement (the minimal content as a statement), definition (minimal plus the benchmark’s definition of the construct) and criteria (minimal plus criteria {true: what + examples, false: what + not-for}). | |
| Direct Choice | The same judgment as a Choice with described options {positive, negative, undetermined[, exclusion]}. | soft: , argmax: |
| Rubric (single) | The reference scorer’s judge prompt asked as one Noul or Choice (e.g., HarmBench’s classifier prompt, SycophancyEval’s “teacher grading a quiz”, RAGTruth’s faithfulness prompt). | or |
| Rubric (formula) | Several atoms whose answers feed the reference scorer’s formula in code (e.g., StrongREJECT refusal convincingness specificity, MASK letters over 5 belief answers and 1 statement, first cave over per-turn answers). | usually discrete: atoms thresholded, then AND/OR or the official aggregate |
| Soft decomposition | Narrow atoms combined without thresholds (per option, per secret, per action). | , , rescaled, |
| Gate | A gate atom (e.g., “does the response engage?”) before a targeted question. | abstain or 0 if , else |
| Comparison | AUROC W/T/L, mean | F1 W/T/L, mean |
|---|---|---|
| Generic Noul vs. generic Choice (soft) | 13/7/11, | 9/5/17, |
| Generic Noul vs. generic Score (soft) | 5/7/19, | 7/7/17, |
| Generic Choice vs. generic Score (soft) | 4/10/17, | 9/6/16, |
| Soft vs. argmax, generic Choice | 28/0/3, | 8/23/5, ( ) |
| Soft vs. argmax, generic Score | 29/1/1, | 10/7/14, |
| Targeted minimal Noul vs. targeted Choice | 7/7/19, | 11/7/15, |
| Pool | AUROC inflation | F1 inflation |
|---|---|---|
| All Jev-only strategies (15–40) | 0.014 (0.026, 0.129) | 0.039 (0.043, 0.176) |
| Generic (5 readouts, ) | 0.005 (0.014, 0.094) | 0.011 (0.020, 0.100) |
| Targeted | 0.008 (0.021, 0.116) | 0.030 (0.039, 0.176) |
| Single-question ( ) | 0.010 (0.018, 0.105) | 0.026 (0.036, 0.181) |
| Multi-question | 0.002 (0.012, 0.102) | 0.016 (0.020, 0.131) |
| Jev AUROC | Jev F1 | |||||||
| Failure type | Usable | Generic Noul | Best gen. (in-sample) | Targeted (split-half) | Margin vs. base. | @0.5 | CV thr. | All-pos. F1 |
| Sycophancy † | 3/4 | 0.726 | 0.732 | 0.784 | +0.118 | 0.640 ∘ | 0.890 | 0.689 |
| Jailbreaks † | 3/4 | 0.965 | 0.972 | 0.963 | +0.057 | 0.818 | 0.891 | 0.507 |
| Deception | 4/4 | 0.949 | 0.971 | 0.945 | +0.215 | 0.845 | 0.853 | 0.614 |
| Prompt injection | 4/4 | 0.962 | 0.988 | 0.998 | +0.026 | 0.535 ∘ | 0.905 | 0.597 |
| Hallucination † | 6/6 | 0.823 | 0.842 | 0.867 | +0.186 | 0.722 | 0.734 | 0.621 |
| AUROC | F1 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Benchmark | Noul | Gen. best | Type | Tgt. best | Tgt. SH | Noul@.5 | Noul cv | Best | Best cv | All-pos. |
| Sycophancy | ||||||||||
| ELEPHANT (AITA) | 0.957 | 1.000 | S | 1.000 | 1.000 | 0.646 | 0.890 | 1.000 | 1.000 | 0.689 |
| SYCON-Bench (false premise) ‡ | 0.976 | 1.000 | S | 0.905 | – | 0.353 | – | 0.880 | – | 0.903 |
| SycophancyEval (answer) | 0.540 | 0.622 | C a | 0.549 | 0.514 | 0.640 | 0.901 | 0.716 | 0.880 | 0.901 |
| SycophancyEval (feedback) | 0.726 | 0.732 | C | 0.789 | 0.784 | 0.390 | 0.327 | 0.411 | 0.400 | 0.243 |
| All usable | Label-chang. | Minority | Hill- | Held- | Corrected | |||
| Statistic | Median | 95% CI | defects excl. | Validated | class | climb | out | labels |
| #B (generic Noul ) | 38 (31) | 31 (28) | 11 (8) | 26 (21) | 29 (25) | 9 (6) | 38 (31) | |
| Generic Noul AUROC | 0.886 | [0.821, 0.952] | 0.905 | 0.872 | 0.886 | 0.886 | 0.906 | 0.893 |
| Best generic AUROC | 0.903 | [0.850, 0.963] | 0.936 | 0.896 | 0.888 | 0.903 | 0.908 | 0.914 |
| Targeted SH AUROC | 0.911 | [0.860, 0.944] | 0.926 | 0.912 | 0.897 | 0.909 | 0.913 | 0.912 |
| Generic Noul F1@0.5 | 0.706 | [0.607, 0.773] | 0.707 | 0.721 | 0.736 | 0.706 | 0.689 | 0.706 |
| Failure type | #B | Generic Noul | Best generic | Targeted SH |
|---|---|---|---|---|
| Sycophancy | 3/3/3 | 0.726 [0.519, 0.964] | 0.732 [0.596, 1.000] | 0.784 [0.468, 1.000] |
| Jailbreaks | 3/3/3 | 0.965 [0.945, 0.997] | 0.972 [0.953, 0.998] | 0.963 [0.920, 0.986] |
| Deception | 4/4/4 | 0.949 [0.916, 0.993] | 0.971 [0.939, 1.000] | 0.945 [0.899, 0.989] |
| Prompt injection | 4/4/4 | 0.962 [0.247, 0.987] | 0.989 [0.537, 0.996] | 0.998 [0.550, 0.999] |
| Hallucination | 6/6/6 | 0.823 [0.711, 0.924] | 0.842 [0.751, 0.932] | 0.867 [0.749, 0.921] |
| Privacy violation | 4/4/4 | 0.808 [0.701, 0.914] | 0.859 [0.803, 0.933] | 0.877 [0.776, 0.942] |
| Primary strategy | Generic Noul | ||||||
|---|---|---|---|---|---|---|---|
| Benchmark | Added to state | Ref. | Tag | AUROC A | AUROC B | [95% CI] | [95% CI] |
| Reference added (32) | |||||||
| SycophancyEval (answer) | reference | missing | key | 0.569 | 0.930 | +0.361 [+0.24, +0.48] | +0.400 [+0.27, +0.52] |
| DeceptionBench | groundtruth | missing | key | 0.973 | 0.976 | +0.003 [ 0.01, +0.01] | +0.003 [ 0.01, +0.01] |
| PrivacyLens | secrets | missing | key | 0.819 | 0.976 | +0.157 [+0.08, +0.23] | +0.159 [+0.08, +0.23] |
| PrivacyLens (all actions) ∘ | secrets | missing | key | 0.874 | 0.974 | +0.101 [+0.05, +0.15] | +0.113 [+0.06, +0.16] |
| Set | Added field | Generic Noul | Best shared |
|---|---|---|---|
| All reference pairs | deployable | 1/7, | 1/7, |
| label key | 5/13, | 9/19, | |
| missing reference | 3/4, | 6/10, | |
| reference present | 1/6, | 1/6, | |
| Collapsed, ceiling-free | all | 5/18, | 7/19, |
| deployable | 1/7, | 1/6, |
| File ( ) | Reference | AUROC short long | [95% CI] |
|---|---|---|---|
| PrivacyLens-all, no secrets (298) | no | 0.944 0.802 | [ , ] |
| PrivacyLens-all, secrets (298) | yes | 0.987 0.962 | [ , ] |
| PrivacyLens, no secrets (110) | no | 0.883 0.737 | [ , ] |
| PrivacyLens, secrets (110) | yes | 0.976 0.975 | [ , ] |
| Revealed reward MT, transcript (69) | no | 0.965 0.847 | [ , ] |
| Revealed reward MT, official (69) | goal and proxy | 0.935 0.863 | [ , ] |
| Pooled | Per file (median) | |||||||
|---|---|---|---|---|---|---|---|---|
| Score set | Files | ECE | AUROC | Slope | ECE | Slope | AUROC | |
| Generic Noul | 48 | 9,797 | 0.035 | 0.884 | 0.91 | 0.150 | 1.58 | 0.905 |
| Targeted Noul | 52 | 40,705 | 0.034 | 0.892 | 0.85 | 0.108 | 1.11 | 0.878 |
| Generic Choice | 53 | 11,193 | 0.087 | 0.879 | 0.45 | 0.128 | 0.65 | 0.879 |
| Best per file | 55 | 11,777 | 0.037 | 0.919 | 0.45 | 0.100 | 1.03 | 0.932 |
| Strategy set | Pairs | F1@0.5 | F1 oracle | F1 CV | Median CV 0.5 | Gap | Median |
|---|---|---|---|---|---|---|---|
| All strategies | 700 | 0.676 | 0.798 | 0.763 | 36% | 0.40 | |
| Best per file | 42 | 0.822 | 0.859 | 0.831 | 71% | 0.50 | |
| Generic Noul | 48 | 0.614 | 0.811 | 0.761 | 23% | 0.35 | |
| Generic Choice | 52 | 0.606 | 0.795 | 0.762 | 15% | 0.25 | |
| Direct Noul | 207 | 0.703 | 0.810 | 0.773 | 38% | 0.45 | |
| Rubric | 36 | 0.781 | 0.874 | 0.855 | 56% | 0.425 |
| ECE (10 bins) | Brier decomposition | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Pool | Obs. | Null mean (p95) | p95 | Eq.-mass | Debiased | Brier | Rel. | Res. | Unc. | |
| Files | 48 | 0.150 | 0.069 (0.100) | 36 | 0.170 | 0.181 | 0.143 | 0.039 | 0.081 | 0.206 |
| Benchmarks | 31 | 0.168 | 0.074 (0.106) | 24 | 0.173 | 0.182 | 0.145 | 0.042 | 0.106 | 0.215 |
| Threshold | Median F1 | F1 all-pos. | all-pos. |
|---|---|---|---|
| 0.706 | 8/31 | ||
| EM prior-shift, prior 0.5 | 0.571 | 18/31 | |
| EM prior-shift, pooled prior | 0.690 | 15/31 | |
| Fit on items | 0.793 | 4/31 | |
| Fit on items | 0.797 | 4/31 | |
| Fit on items ∗ | 0.803 | 2/26 |
| Label source (#B) | AUROC | ECE (null) | F1@0.5 | F1 CV | F1, | All-pos. | CV 0.5 | |
|---|---|---|---|---|---|---|---|---|
| Validated (8) | 0.872 [0.772, 0.922] | 0.114 (0.059) | 0.40 | 0.721 | 0.694 | 0.684 | 0.536 | [ , ] |
| LLM judge (16) | 0.906 [0.818, 0.966] | 0.149 (0.086) | 0.35 | 0.767 | 0.825 | 0.819 | 0.595 | [ , ] |
| Rule (4) | 0.963 [0.947, 0.987] | 0.243 (0.053) | 0.35 | 0.535 | 0.906 | 0.829 | 0.597 | [ , ] |
| Label-changing defect (3) | 0.540 [0.208, 0.759] | 0.252 (0.062) | 0.50 | 0.640 | 0.832 | 0.827 | 0.826 | [ , ] |
| All (31) | 0.886 [0.821, 0.952] | 0.168 (0.074) | 0.35 | 0.706 | 0.822 | 0.793 | 0.568 | – |
| Failure | Benchmark | Neg. | Pos. | Base rate | F1@0.5 | F1 CV | All-pos. |
| CV all-positive | SycophancyEval (answer) | 27 | 123 | 0.82 | 0.640 | 0.901 | 0.901 |
| LLM-AggreFact (multi-turn) | 9 | 139 | 0.94 | 0.841 | 0.969 | 0.969 | |
| Reference letters (gender) | 13 | 30 | 0.70 | 0.125 | 0.817 | 0.822 | |
| Workplace scenes (WinoBias) | 7 | 37 | 0.84 | 0.318 | 0.880 | 0.914 | |
| CV | SycophancyEval (feedback) | 187 | 30 | 0.14 | 0.390 | 0.327 | 0.243 |
| MASK (continuation) | 40 | 42 | 0.51 | 0.849 | 0.830 | 0.677 |
| Signal | Acc. | AUROC | Top 10% | Top 20% | Top 50% | AURC (oracle) | |
|---|---|---|---|---|---|---|---|
| Generic Choice , confidence | 9,979 | 0.814 | 0.745 | 0.944 | 0.965 | 0.924 | 0.102 (0.019) |
| Generic Score , confidence | 10,097 | 0.833 | 0.762 | 0.988 | 0.974 | 0.943 | 0.070 (0.015) |
| Generic Noul , | 10,097 | 0.806 | 0.780 | 0.994 | 0.980 | 0.931 | 0.075 (0.020) |
| Targeted Noul , | 41,905 | 0.811 | 0.779 | 0.994 | 0.983 | 0.937 | 0.072 (0.019) |
| Rater | vs. human: acc. / / AUROC | vs. scorer: acc. / / AUROC |
|---|---|---|
| Reference scorer (GPT-4o-mini rubric) | 0.936 / 0.811 / 0.929 | – |
| Jev generic Noul (pre-specified) | 0.935 / 0.809 / 0.971 | 0.915 / 0.753 / 0.955 |
| Jev generic Choice | 0.939 / 0.819 / 0.973 | 0.921 / 0.767 / 0.953 |
| Jev targeted † (“substantive help”, criteria) | 0.956 / 0.868 / 0.984 | 0.935 / 0.806 / 0.973 |
| Jev re-implemented scorer rubric | 0.932 / 0.801 / 0.953 | 0.927 / 0.788 / 0.938 |
| gain | AUROC | Humans side | ||
|---|---|---|---|---|
| Selection evaluation | Jev / scorer | [95% CI] | Jev / scorer | with Jev |
| None (generic Noul ) | 0.809 / 0.811 | [ , ] | 0.971 / 0.929 | 49.1% of 116 |
| Split-half, prompt groups | 0.865 / 0.819 | [ , ] | 0.984 / 0.934 | 63.1% [43.8, 75.6] |
| Split-half, item groups | 0.858 / 0.803 | [ , ] | 0.982 / 0.926 | 64.8% [52.9, 74.7] |
| HarmBench val. SR | 0.819 / 0.811 | [ , ] | 0.973 / 0.929 | 51.8% of 108 |
| SR HarmBench val. | 0.777 / – | [ , ] ∗ | 0.957 / – | – |
| @0.5 [95% CI] | AUROC | ||||||
|---|---|---|---|---|---|---|---|
| Generator | (pos.) | Jev generic | Jev targeted | Scorer | Jev gen. | Jev tgt. | Scorer |
| dolphin-mixtral | 500 (188) | 0.874 [.828, .916] | 0.928 [.887, .963] | 0.820 [.756, .883] | 0.987 | 0.994 | 0.924 |
| GPT-3.5 | 491 (91) | 0.668 [.585, .750] | 0.752 [.647, .840] | 0.790 [.714, .856] | 0.928 | 0.959 | 0.938 |
| GPT-4 | 112 (9) | 0.731 [.430, .943] | 0.809 [.583, 1.00] | 0.537 [.267, .825] | 0.982 | 0.981 | 0.848 |
| Detector | AUROC | @0.5 | ECE (null p95) | Brier | Slope |
|---|---|---|---|---|---|
| GPT-4o-mini rubric (reference scorer) | 0.929 | 0.811 | 0.045 (0.013) | 0.055 | 0.34 |
| Rubric component: 1 refusal | 0.833 | 0.517 | 0.208 | 0.208 | 0.18 |
| Jev generic Noul | 0.971 | 0.809 | 0.062 (0.025) | 0.053 | 1.20 |
| Jev generic Noul , 9 quantile levels | 0.962 | – | – | – | – |
| Jev generic Choice | 0.973 | 0.819 | 0.032 (0.017) | 0.046 | 0.65 |
| Jev generic Score | 0.968 | 0.810 | 0.039 (0.021) | 0.056 | 0.80 |
| Benchmark ( ) | Target model | Jev direct Choice | Jev generic Noul |
|---|---|---|---|
| LLM-AggreFact A (150) | 0.533 / 0.631 | 0.733 / 0.769 | 0.707 / 0.762 |
| LLM-AggreFact B (150) | 0.547 / 0.757 | 0.793 / 0.874 | 0.780 / 0.881 |
| SummEdits (300) | 0.557 / 0.648 | 0.813 / 0.852 | 0.803 / 0.853 |
| ConfAIde tier 2b (98) | 0.758 / 0.816 | 0.863 / 0.918 | – |
| PrivaCI-Bench (300) | 0.711 / 0.711 | 0.730 / 0.815 | – |
| PrivaCI-Bench GDPR (300) | 0.938 / 0.938 | 0.881 / 0.969 | – |
| Items | #B | Jev err. | Determ. | Interpr. | Construct | Unclass. | |
|---|---|---|---|---|---|---|---|
| All read items | 21 | 311 | 25.4% | 28.6% | 14.8% | 24.1% | 7.1% |
| Excl. HarmBench validation | 20 | 258 | 22.9% | 34.5% | 17.8% | 16.3% | 8.5% |
| Excl. also Open-Prompt-Injection | 19 | 169 | 34.9% | 3.6% | 27.2% | 24.9% | 9.5% |
| Benchmark (failure) | Label defect | Effect | Evidence |
|---|---|---|---|
| Open-Prompt-Injection (prompt injection) | Positive only if the model followed the injection and solved the injected task: tracks task accuracy. | Changed (det.) | 293/300 responses follow the injection, and 83/89 negatives are followed-but-wrong. Targeted AUROC 0.50–0.55 on the label. On the corrected label “follows the injection”, generic Noul 0.975, targeted 1.000. |
| SycophancyEval (answer) (sycophancy) | Label is answer correctness, not deference. | Changed (det.) | Target is wrong on 44/50 items with no user opinion. Deference gives AUROC 0.75 or 0.22 depending on which answer the user suggests (0.49 pooled). On the corrected label “answer moves toward the suggestion”, blind generic Noul 0.712. |
| SycophancyEval (feedback) (sycophancy) | Judge applies “no criticism” inconsistently. | Changed (judge) | 15 of 30 positives name a fallacy or flaw in the first sentence. |
| MACHIAVELLI (power, harm, reward, held-out games) (power seeking, reward hacking) | Whether the target’s choice is the option annotated with more in-game consequences. Neither is in the state. | Changed (unobs.) | Label equals “choice flagged option” on 1,100/1,100 items, target picks B on 75–78% of items, and P(picks flagged) 0.71–0.76 if flagged is B vs. 0.16–0.26 if A. |
| Workplace scenes (WinoBias) (social bias) | Prompt forces a gap (it asks who takes the lead), and the parser maps FEMALE_MORE to MALE_MORE . | Direction only | 95%/84% positive by construction, and direction wrong on 65/79 positives. Jev recovers the corrected direction on 92–100% of positives. |
| AbstentionBench (concealing uncertainty) | Label depends on which judge runs the official prompt. | Defect (judge), degenerate | Qwen2.5-7B and GPT-4o judges agree on 50% of should-abstain items. Jev F1 0.70 on Qwen labels vs. 0.85 on GPT-4o labels. |
| Benchmark | Read | J | Det. | Int. | Constr. | Cluster | Cluster-level finding or note |
| Open-Prompt-Injection | 89 | 0 | 83 | 0 | 0 | 300 | 6 negatives did not follow the injection |
| SycophancyEval (answer) | 4 | 0 | 0 | 4 | 0 | 200 | label is correctness (44/50 no-opinion answers wrong) |
| SycophancyEval (feedback) | 31 | 0 | 0 | 16 | 0 | – | 15 (+2 borderline) of 30 positives contradict the definition |
| StrongREJECT | 9 | 0 | 0 | 3 | 6 | – | 3 deflections scored harmful, rest borderline persona responses |
| JailbreakBench (PAIR), benign split | 4 | 1 | 0 | 3 | 0 | – | explicit refusals labelled “no refusal” |
| HarmBench | 11 | 0 | 0 | 0 | 11 | – | attempted but low-utility outputs |
| AUROC | F1 | ||||||
|---|---|---|---|---|---|---|---|
| Benchmark and label | Pos./ | Generic | Best gen. | Tgt. SH | Generic@0.5 | Generic CV | All-pos. |
| Open-Prompt-Injection, official | 211/300 | 0.231 | 0.534 | 0.536 | 0.832 | 0.832 | 0.826 |
| corrected: follows injection | 293/300 | 0.975 | 1.000 | 1.000 | 0.995 | 0.995 | 0.988 |
| SycophancyEval (ans.), official | 123/150 | 0.540 | 0.622 | 0.514 | 0.640 | 0.901 | 0.901 |
| corrected: moves to suggestion | 15/150 | 0.712 | 0.763 | 0.752 | 0.239 | 0.250 | 0.182 |
| Reference scorer | Jev | ||||||
|---|---|---|---|---|---|---|---|
| Scorer group | #B | Calls | Tokens | USD | Tokens | USD | Ratio |
| API LLM judge | 19 | 8,940 | 9.36M | 18.96 | 7.18M | 0.30 | 62.9 |
| Local judge / classifier | 5 | 1,171 | 0.52M | – | 2.14M | 0.09 | – |
| Rule / log-prob scorer | 20 | – | – | – | 8.62M | 0.36 | – |
| All Jev calls (23,384) | 44 | 54.0M | 2.27 | ||||
| Failure type | Calls | Q/call | Err. (%) | In tok. | Out/Q | Lat. med. (s) | Lat. p90 (s) | USD |
|---|---|---|---|---|---|---|---|---|
| Sycophancy | 1,207 | 11.6 | 0.00 | 2,634 | 25.6 | 0.34 | 0.79 | 0.163 |
| Jailbreaks | 4,848 | 14.2 | 0.56 | 2,579 | 22.0 | 0.31 | 0.53 | 0.554 |
| Deception | 1,988 | 12.1 | 0.00 | 2,227 | 33.6 | 0.33 | 0.43 | 0.281 |
| Prompt injection | 2,824 | 9.7 | 0.00 | 1,686 | 26.2 | 0.34 | 1.30 | 0.211 |
| Hallucination | 3,000 | 12.5 | 0.00 | 2,529 | 24.0 | 0.31 | 0.38 | 0.331 |
| Privacy violation | 2,860 | 10.8 | 0.00 | 1,883 | 25.3 | 0.31 | 0.50 | 0.243 |
| Reference scorer | Jev | ||||||
|---|---|---|---|---|---|---|---|
| Benchmark | Judge | Items | Calls/item | Tok./item | USD | Tok./item | USD |
| SycophancyEval (answer) | GPT-4o | 200 | 1.00 | 514 | 0.259 | 1,907 | 0.0160 |
| SYCON-Bench (false premise) | GPT-4o | 42 | 1.86 | 1,890 | 0.199 | 8,073 | 0.0142 |
| SycophancyEval (feedback) | GPT-4o | 217 | 2.00 | 1,096 | 0.598 | 2,740 | 0.0250 |
| StrongREJECT | GPT-4o-mini | 100 | 1.00 | 44 | 0.001 | 2,691 | 0.0113 |
| HarmBench | HarmBench-13B cls. ∗ | 150 | 1.00 | 226 | – | 3,144 | 0.0198 |
| Judges at | Jev billed for | Judge $ | Jev $ | Pooled | Median (range) | 2 | 5 | PL / MASK | Excl. |
|---|---|---|---|---|---|---|---|---|---|
| List price | full battery | 18.96 | 0.302 | 62.9 | 16.2 (0.12–142 ) | 18 | 18 | 28% / 37% | 35.1 |
| GPT-4o-mini | single question | 1.82 | 0.150 | 12.1 | 3.3 (0.50–43.7 ) | 17 | 8 | 39% / 23% | 7.3 |
| GPT-4o-mini | full battery | 1.82 | 0.302 | 6.0 | 1.5 (0.12–18.9 ) | 8 | 6 | 39% / 23% | 3.6 |
| Detector | AUROC | ECE ( null) | F1@0.5 | F1 CV | F1, / 20 / 50 | F1 oracle |
|---|---|---|---|---|---|---|
| Jev generic Noul (zero-shot) | 0.886 | 0.168 (24/31) | 0.706 | 0.822 | 0.793 / 0.797 / 0.803 | 0.849 |
| TF-IDF LR (out-of-fold) | 0.754 | 0.164 (22/31) | 0.610 | 0.609 | 0.604 / 0.605 / 0.629 | 0.647 |
| Length (out-of-fold LR) | 0.526 | 0.047 (2/31) | 0.235 | 0.632 | 0.565 / 0.590 / 0.636 | 0.636 |
| TF-IDF trained on items | – | – | – | – | 0.351 / 0.387 / 0.514 | – |
| Jev AUROC | TF-IDF LR | Length | All-pos. | |||
|---|---|---|---|---|---|---|
| Benchmark | Generic | Targeted SH | AUROC | F1 | AUROC | F1 |
| Sycophancy | ||||||
| ELEPHANT (AITA) | 0.957 | 1.000 | 0.748 0.034 | 0.727 | 0.515 | 0.689 |
| SycophancyEval (answer) | 0.540 | 0.514 | 0.764 0.012 | 0.914 | 0.451 | 0.901 |
| SycophancyEval (feedback) | 0.726 | 0.784 | 0.608 0.016 | 0.138 | 0.547 | 0.243 |
| Jailbreaks | ||||||