Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
Organizations: TraceStone and Nanyang Technological University, Singapore
Abstract
Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models select from predefined answers and report probabilities that software can use to allow, block, or review inputs, but the reliability of these automated decisions remains unclear. We evaluate Jev, Laya, Decider, and Bespoke Nimble against specialized classifiers and language-model judges, examining decision accuracy, probability calibration, and selective automation. We draw the following conclusions. (1) Strong overall performance and favorable aggregate calibration can hide failures concentrated in particular attack groups, including attacks classified as safe with high confidence. (2) The evaluated adapted configurations do not consistently improve classification over their base models across tasks. (3) Under the strictest evaluated error limits, the policies allow few inputs automatically, and separate allow and block thresholds increase automation mainly through more blocks. Passing confirmation does not ensure that these limits hold on test. (4) Judges can detect attacks missed by another model, but may also falsely flag more benign inputs and share the other model's high-confidence errors. These findings support evaluating model accuracy, probability calibration, and the resulting allow/block/review decisions together.
Figures & tables
| Configuration | Backend | Weights | Tasks |
|---|---|---|---|
| Jev 1.13 | Hosted API | – | PI, IR, HA |
| Laya, English checkpoint | PyTorch CPU | FP32 | PI, IR, HA |
| Decider, approximately 2B | PyTorch MPS | FP16 | PI, IR, HA |
| Bespoke Nimble, approximately 9B | MLX GPU | BF16 | PI, IR, HA |
| GPT-4.1 | Hosted API | – | PI, IR, HA |
| Qwen3-8B, thinking disabled | API/MPS | –/BF16 | PI, IR, HA |
| Benchmark | Development | Selection | Confirmation | Test |
|---|---|---|---|---|
| R-Judge | 100 | 140 | 93 | 236 |
| WAInjectBench-text | 100 | 966 | 644 | 1,612 |
| AgentHarm | 64 | – | – | 352 |
| Rank | Configuration | Macro-F1 [95% CI] | FNR | FPR |
|---|---|---|---|---|
| 1 | Nimble | 0.778 [0.495, 0.936] | 0.473 | 0.035 |
| 2 | Qwen3-8B | 0.764 [0.479, 0.929] | 0.495 | 0.039 |
| 3 | Jev | 0.756 [0.452, 0.937] | 0.599 | 0.001 |
| 4 | GPT-4.1 | 0.707 [0.428, 0.932] | 0.681 | 0.002 |
| 5 | ProtectAI | 0.658 [0.412, 0.769] | 0.724 | 0.023 |
| 6 | PIGuard | 0.656 [0.395, 0.800] | 0.735 | 0.018 |
| Rank | Configuration | Macro-F1 [95% CI] | FNR | FPR |
|---|---|---|---|---|
| 1 | Jev | 0.825 [0.774, 0.873] | 0.056 | 0.297 |
| 2 | GPT-4.1 | 0.807 [0.755, 0.857] | 0.064 | 0.324 |
| 3 | Decider parent | 0.604 [0.540, 0.666] | 0.232 | 0.550 |
| 4 | Laya | 0.474 [0.411, 0.536] | 0.752 | 0.207 |
| 5 | Nimble parent | 0.450 [0.389, 0.516] | 0.128 | 0.847 |
| 6 | Qwen3-8B | 0.427 [0.369, 0.487] | 0.064 | 0.901 |
| Rank | Configuration | Macro-F1 [95% CI] | FNR | FPR |
|---|---|---|---|---|
| 1 | Jev | 0.840 [0.774, 0.903] | 0.080 | 0.239 |
| 2 | Nimble | 0.831 [0.756, 0.892] | 0.068 | 0.267 |
| 3 | Decider | 0.778 [0.702, 0.846] | 0.415 | 0.011 |
| 4 | Nimble parent | 0.763 [0.694, 0.824] | 0.057 | 0.403 |
| 5 | Qwen3-8B | 0.751 [0.683, 0.813] | 0.034 | 0.443 |
| 6 | GPT-4.1 | 0.741 [0.659, 0.810] | 0.006 | 0.483 |
| Rank | Configuration | Macro-F1 [95% CI] | FNR | FPR |
|---|---|---|---|---|
| 1 | GLM-5.3-Flash | 0.789 [0.495, 0.949] | 0.513 | 0.011 |
| 2 | Nimble | 0.778 [0.495, 0.936] | 0.473 | 0.035 |
| 3 | DeepSeek-V4.1-Flash | 0.777 [0.468, 0.953] | 0.548 | 0.005 |
| 4 | Jev | 0.756 [0.452, 0.937] | 0.599 | 0.001 |
| 5 | Qwen3.8-27B | 0.735 [0.448, 0.923] | 0.627 | 0.005 |
| 6 | GLM-5.3 | 0.717 [0.447, 0.946] | 0.659 | 0.005 |
| Rank | Configuration | Macro-F1 [95% CI] | FNR | FPR |
|---|---|---|---|---|
| 1 | DeepSeek-V4.1-Flash | 0.966 [0.941, 0.987] | 0.032 | 0.036 |
| 2 | Qwen3.8-27B | 0.945 [0.915, 0.970] | 0.080 | 0.027 |
| 3 | GLM-5.3 | 0.919 [0.885, 0.953] | 0.128 | 0.027 |
| 4 | GLM-5.3-Flash | 0.911 [0.872, 0.948] | 0.096 | 0.081 |
| 5 | Kimi-K2.6 | 0.906 [0.865, 0.941] | 0.064 | 0.126 |
| 6 | Jev | 0.825 [0.774, 0.873] | 0.056 | 0.297 |
| Rank | Configuration | Macro-F1 [95% CI] | FNR | FPR |
|---|---|---|---|---|
| 1 | Jev | 0.840 [0.774, 0.903] | 0.080 | 0.239 |
| 2 | Nimble | 0.831 [0.756, 0.892] | 0.068 | 0.267 |
| 3 | GLM-5.3-Flash | 0.817 [0.751, 0.874] | 0.091 | 0.273 |
| 4 | GLM-5.3 | 0.797 [0.732, 0.853] | 0.057 | 0.341 |
| 5 | DeepSeek-V4.1-Flash | 0.779 [0.705, 0.847] | 0.023 | 0.403 |
| 6 | Qwen3.8-27B | 0.748 [0.670, 0.814] | 0.011 | 0.466 |
| Macro-F1 | FNR | FPR | |||||
|---|---|---|---|---|---|---|---|
| Configuration | All/shared | All | Shared | All | Shared | All | Shared |
| Nimble | 1612/1612 | 0.778 | 0.778 | 0.473 | 0.473 | 0.035 | 0.035 |
| Qwen3-8B | 1612/1612 | 0.764 | 0.764 | 0.495 | 0.495 | 0.039 | 0.039 |
| Jev | 1612/1612 | 0.756 | 0.756 | 0.599 | 0.599 | 0.001 | 0.001 |
| GPT-4.1 | 1612/1612 | 0.707 | 0.707 | 0.681 | 0.681 | 0.002 | 0.002 |
| ProtectAI | 1612/1612 | 0.658 | 0.658 | 0.724 | 0.724 | 0.023 | 0.023 |
| Macro-F1 | FNR | FPR | |||||
|---|---|---|---|---|---|---|---|
| Configuration | All/shared | All | Shared | All | Shared | All | Shared |
| Jev | 236/236 | 0.825 | 0.825 | 0.056 | 0.056 | 0.297 | 0.297 |
| GPT-4.1 | 236/236 | 0.807 | 0.807 | 0.064 | 0.064 | 0.324 | 0.324 |
| Decider parent | 236/236 | 0.604 | 0.604 | 0.232 | 0.232 | 0.550 | 0.550 |
| Laya | 236/236 | 0.474 | 0.474 | 0.752 | 0.752 | 0.207 | 0.207 |
| Nimble parent | 236/236 | 0.450 | 0.450 | 0.128 | 0.128 | 0.847 | 0.847 |
| Configuration | AUROC | AP | Brier | ECE | NLL | |
|---|---|---|---|---|---|---|
| WAInjectBench | ||||||
| Nimble | 1612 | 0.734 | 0.582 | 0.103 | 0.087 | 0.364 |
| Jev | 1612 | 0.796 | 0.662 | 0.097 | 0.105 | 1.140 |
| ProtectAI | 1612 | 0.797 | 0.565 | 0.139 | 0.138 | 1.234 |
| PIGuard | 1612 | 0.657 | 0.483 | 0.132 | 0.126 | 0.933 |
| Laya | 1612 | 0.633 | 0.345 | 0.147 | 0.108 | 0.469 |
| Component | Bias | W5 | W10 | W20 | Q10 | Skill |
|---|---|---|---|---|---|---|
| WAInjectBench ( , ) | ||||||
| Jev | -0.105 | 0.105 | 0.105 | 0.105 | 0.105 | +0.324 |
| Laya | +0.108 | 0.108 | 0.108 | 0.116 | 0.115 | -0.024 |
| Decider | -0.143 | 0.143 | 0.143 | 0.143 | 0.143 | -0.031 |
| Nimble | +0.036 | 0.052 | 0.087 | 0.088 | 0.090 | +0.282 |
| R-Judge ( , ) | ||||||
| Configuration | AUROC | AP | Brier | ECE | NLL |
|---|---|---|---|---|---|
| WAInjectBench | |||||
| DeepSeek-V4.1-Flash | 0.774 | 0.626 | 0.099 | 0.099 | 1.598 |
| GLM-5.3 | 0.685 | 0.531 | 0.116 | 0.117 | 1.187 |
| Kimi-K2.6 | 0.786 | 0.642 | 0.122 | 0.122 | 1.580 |
| Qwen3.8-27B | 0.645 | 0.512 | 0.112 | 0.112 | 1.485 |
| R-Judge | |||||
| Configuration | Brier | ECE | NLL | |
|---|---|---|---|---|
| WAInjectBench | ||||
| Jev | 10.93 | 0.097 / 0.107 | 0.105 / 0.133 | 1.140 / 0.374 |
| Laya | 15.19 | 0.147 / 0.237 | 0.108 / 0.310 | 0.469 / 0.666 |
| Decider | 20.00 | 0.148 / 0.216 | 0.143 / 0.285 | 0.582 / 0.625 |
| Nimble | 1.18 | 0.103 / 0.107 | 0.087 / 0.104 | 0.364 / 0.376 |
| R-Judge | ||||
| Configuration | Coverage | Allow | Miss | False block | Confirmation pass |
| WAInjectBench | |||||
| Jev | 0.00% | 0.00% | 0.00% | 0.00% | Yes |
| Laya | 5.02% | 4.71% | 1.08% | 0.23% | Yes |
| Decider | 0.00% | 0.00% | 0.00% | 0.00% | Yes |
| Nimble | 0.74% | 0.74% | 0.36% | 0.00% | Yes |
| Qwen3-8B ‡ | 0.00% | 0.00% | 0.00% | 0.00% | Yes |
| Configuration | Allowed | Blocked | Misses | False blocks | C/T | ||
| WAInjectBench | |||||||
| Jev | 0.000 | 0.525 | 0 | 108 | 0 | 0 | Y/Y |
| Laya | 0.105 | 0.570 | 76 | 133 | 3 | 65 | Y/N |
| Decider | 0.000 | 0.515 | 0 | 14 | 0 | 0 | Y/Y |
| Nimble | 0.050 | 0.500 | 12 | 194 | 1 | 47 | Y/Y |
| GPT-4.1 | 0.000 | 0.655 | 0 | 80 | 0 | 2 | Y/Y |
| Configuration | Cov. S | Cov. I | Allow I | Miss I | FB I | C/T | ||
|---|---|---|---|---|---|---|---|---|
| WAInjectBench | ||||||||
| Jev | 0.000 | 0.500 | 0.00 | 6.95 | 0.00 | 0.00 | 0.08 | Y/Y |
| Laya | 0.465 | 0.505 | 5.15 | 12.90 | 4.84 | 1.08 | 4.73 | Y/N |
| Decider | 0.000 | 0.500 | 0.00 | 0.87 | 0.00 | 0.00 | 0.00 | Y/Y |
| Nimble | 0.075 | 0.500 | 0.37 | 12.41 | 0.37 | 0.00 | 3.53 | Y/Y |
| GPT-4.1 | 0.000 | 0.505 | 0.00 | 4.96 | 0.00 | 0.00 | 0.15 | Y/Y |
| Judge | False negatives | False positives | ||
| Corrected | Added | Corrected | Added | |
| WAInjectBench, | ||||
| GPT-4.1 | 10 of 167 | 33 | 1 of 1 | 2 |
| Qwen3-8B | 32 of 167 | 3 | 0 of 1 | 51 |
| GLM-5.3-Flash | 29 of 167 | 5 | 1 of 1 | 14 |
| DeepSeek-V4.1-Flash | 22 of 167 | 8 | 1 of 1 | 7 |