TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models
Organizations: DFKI GmbH · RPTU Kaiserslautern · Imperial College London · Mila - Quebec AI Institute
Abstract
System One models output calibrated probabilities over typed answers such as categorical choices, ordinal levels, or binary outcomes, via a non-generative interface. Software can act on these probabilities through thresholds, cost-weighted choices, and escalation rules. Consequently, if these probabilities are miscalibrated or wording-sensitive, the software ma take unintended actions leaving human operators with no textual rationale to inspect. Current evaluations largely report accuracy and calibration on public classification datasets without a clear reference. We present TypedBench, a benchmark for typed decision models like Jev built from seven policy-labelled generators and nine evaluation suites. We report accuracy as median and range across paraphrases, and calibration error relative to the finite-sample noise floor of a matched, perfectly calibrated predictor. We assess probability quality through selective prediction, ordinal proper scoring rules, and realised cost under asymmetric cost matrices. We evaluate a hosted model, an open encoder, and a family of open decoders spanning 0.8B-9B parameters on identical items. The hosted model follows the stated policy but is wording-sensitive and systematically underconfident; under asymmetric costs, using its probabilities can be worse than taking its top answer. Decoders route exactly and are slow as options or questions are added. The decoder is least accurate on policy questions and degrades with more options. Overall, typed decision models must be evaluated jointly on policy adherence, wording robustness, probability quality, and induced decision outcomes.
Figures & tables
| Generator | Questions | Capability isolated |
|---|---|---|
| support tickets | queue (up to 12), anger, priority (5 levels) | applying a three-term policy |
| phishing email | phishing, attack class (5-way), urgency (4 levels), five sub-questions | decomposing a judgement into evidence |
| RAG relevance | best passage among ( ), relevant, relevance (4 levels) | choosing among many options with hard distractors |
| log triage | severity (4 levels), owning team (5), page on-call | a lookup table and a two-field rule in a JSON record |
| policy compliance | compliant, violated clause or none, severity (3 levels) | matching a request to numbered clauses |
| guardrail intent | intent (4), block, risk (4 levels) | separating attack mechanism from harm |
| Asks | Varies | Metric | |
|---|---|---|---|
| A | Do confidences match accuracy? | unknowable and label-noise arms | ECE/null, , AURC |
| B | Is accuracy stable under rewording? | paraphrases, criteria variants, adversarial wordings, option order | accuracy range, JSD, flips |
| C | Does it scale? | options –255; state length 128–30k tokens | accuracy, latency |
| D | Is it robust? | distractors, perturbations, none-correct, prior shift | accuracy change, abstention, AUROC |
| E | Do co-asked questions interfere? | 2–20 co-questions | answer changes, latency |
| F | Are scores ordered? | severity ladders; 3/5/10-level rubrics | monotonicity, Spearman |
| Question | Jev | Laya english | Laya typed-dec. | Kev 0.8B | Kev 4B | Kev 9B |
|---|---|---|---|---|---|---|
| logs.team (lookup) | 1.000 | 0.413 | 0.473 | 0.977 | 0.983 | 1.000 |
| logs.severity (map) | 1.000 | 0.327 | 0.393 | 0.593 | 0.747 | 0.720 |
| logs.page (2-field) | 0.990 | 0.257 | 0.370 | 0.797 | 0.987 | 0.860 |
| policy.compliant | 1.000 | 0.713 | 0.743 | 0.773 | 0.860 | 0.853 |
| policy.clause | 0.993 | 0.387 | 0.393 | 0.693 | 0.987 | 0.903 |
| policy.severity | 0.750 | 0.357 | 0.407 | 0.610 | 0.837 | 0.857 |
| Model | argmax | Bayes | value | 95% interval |
|---|---|---|---|---|
| Jev 1.13.0 | 1,080 | 2,300 | [ , ] | |
| Laya english | 12,740 | 11,840 | [ , ] | |
| Laya typed-dec. | 11,924 | 10,802 | [ , ] | |
| Kev -0.8B | 13,580 | 9,860 | [ , ] | |
| Kev -4B | 8,980 | 5,000 | [ , ] | |
| Kev -9B | 4,780 | 7,960 | [ , ] |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Generator | Questions (type, cardinality) | Label policy and knobs |
|---|---|---|
| support_tickets 2.0.0 | queue (choice 2–12), is_angry (noul), priority (score 5) | each template declares facts (blocked / money at stake / action requested / question only); urgency = 2 if blocked or money, 1 if action, else 0; priority = urgency if angry if gold/enterprise, cap 4; knobs: cardinality, target tokens, distractors, label noise, unknowable, none-correct; surface variation: openers, register, sign-offs, typos |
| phishing_email 1.1.0 | is_phishing (noul), attack_class (choice 5), urgency (score 4), five sub-nouls | lookalike vs legitimate sender domains, template class; urgency from a declared time-pressure class (none / softened request / deadline within days / within 24 h or threat); decomposition rule: phishing if any sub-question true |
| rag_relevance | best_passage (choice ), is_relevant (noul), relevance (score 4) | fact table over four topics (entity, attribute, value); relevance 3 if the passage states the queried attribute of the queried entity, 2 if it shares only the entity or only the attribute, 1 if only the topic, 0 otherwise; best_passage is the grade-3 passage; is_relevant iff grade 3 |
| log_triage | severity (score 4), owning_team (choice 5), page_oncall (noul) | literal level map; lookup table in the state; page iff severity and env = production |
| policy_compliance | compliant (noul), violated_clause (choice 4–6 incl. none), severity (score 3) | 3–5 clauses tagged minor or major; each request violates exactly one clause or none (30% compliant); compliant iff none is violated; violated_clause = that clause or none; severity 0 if compliant, else 1 (minor) or 2 (major) |
| guardrail_intent 1.1.0 | intent (choice 4), block (noul), risk (score 4) | paraphrase clusters per class; intent by mechanism (tie-break stated in the question); risk = 3 if the goal is harmful by any mechanism, 2 if subversion only, 1 benign-sensitive, 0 otherwise |
| Question | paraphrases | criteria variants | adversarial | negations |
|---|---|---|---|---|
| tickets.queue | 5 | 3 | – | – |
| tickets.is_angry | 5 | – | 2 | 2 |
| tickets.priority | 5 | – | 3 | – |
| phish.is_phishing | 5 | – | 2 | 2 |
| phish.attack_class | 3 | 2 | – | – |
| phish.urgency | 3 | – | – | – |
| Step | Complaint | 3 levels | 5 levels | 10 levels |
|---|---|---|---|---|
| 0 | The app shows a small typo in the settings menu. | 0 | 0 | 0 |
| 1 | One report page loads slowly, about ten seconds. | 0 | 1 | 2 |
| 2 | Exports fail for reports longer than 30 days; other features work. | 1 | 2 | 4 |
| 3 | I cannot log in on any device since this morning. | 2 | 3 | 7 |
| 4 | Our whole team is locked out and payroll runs in two hours. | 2 | 4 | 9 |
| Question (primitive) | Model | Acc (median) | Range | ECE/null | AURC | p50 ms | |
|---|---|---|---|---|---|---|---|
| tickets.priority (score, 5) | Jev 1.13.0 | 0.874 | [0.708, 0.946] | 2.2 | 0.66 | 0.023 | 233 |
| Laya english | 0.343 | [0.220, 0.372] | 3.3 | 5.04 | 0.660 | 59 | |
| Laya typed-dec. | 0.383 | [0.347, 0.412] | 1.3 | 0.97 | 0.660 | 59 | |
| Kev -0.8B | 0.282 | [0.236, 0.358] | 2.7 | 0.67 | 0.661 | 62 | |
| Kev -4B | 0.446 | [0.382, 0.528] | 2.8 | 0.50 | 0.320 | 208 | |
| Kev -9B | 0.565 | [0.476, 0.660] | 2.8 | 0.78 | 0.239 | 325 |
| Question | Model | Acc | smooth ECE | Brier | NLL | cov@5% | flip | JSD | RPS | QWK |
|---|---|---|---|---|---|---|---|---|---|---|
| tickets.queue | Jev 1.13.0 | 1.000 | 0.015 | 0.004 | 0.016 | 1.00 | 0.029 | 0.026 | – | – |
| Laya english | 0.854 | 0.194 | 0.319 | 0.704 | 0.15 | 0.114 | 0.110 | – | – | |
| Laya typed-dec. | 0.874 | 0.106 | 0.212 | 0.432 | 0.67 | 0.066 | 0.042 | – | – | |
| Kev -0.8B | 1.000 | 0.128 | 0.038 | 0.146 | 1.00 | 0.045 | 0.079 | – | – | |
| Kev -4B | 1.000 | 0.106 | 0.025 | 0.115 | 1.00 | 0.032 | 0.043 | – | – | |
| Kev -9B | 1.000 | 0.100 | 0.024 | 0.113 | 1.00 | 0.009 | 0.016 | – | – |
| Jev 1.13.0 | Laya english | Laya typed-dec. | Kev -0.8B | Kev -4B | Kev -9B | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Question | ECE/null | AURC | ECE/null | AURC | ECE/null | AURC | ECE/null | AURC | ECE/null | AURC | ECE/null | AURC |
| logs.team (lookup) | – | 0.000 | 2.5 | 0.356 | 3.7 | 0.359 | 4.3 | 0.002 | 2.6 | 0.000 | 2.0 | 0.000 |
| logs.severity (map) | 2.5 | 0.000 | 4.4 | 0.628 | 1.7 | 0.606 | 2.4 | 0.197 | 4.4 | 0.152 | 4.5 | 0.052 |
| logs.page (2-field) | 3.4 | 0.000 | 14.6 | 0.804 | 6.5 | 0.442 | 3.4 | 0.063 | 3.7 | 0.000 | 3.0 | 0.016 |
| policy.compliant | 3.7 | 0.000 | 1.9 | 0.176 | 1.9 | 0.194 | 2.1 | 0.096 | 2.2 | 0.028 | 1.4 | 0.031 |
| policy.clause | 0.9 | 0.000 | 1.7 | 0.636 | 2.0 | 0.605 | 2.0 | 0.175 | 2.0 | 0.001 | 1.2 | 0.006 |
| Largest accuracy change | Out of scope | Prior shift (0.2 / 0.5 / 0.8) | |||||
|---|---|---|---|---|---|---|---|
| Model | perturbation | distractors | abstain | AUROC | conf. out / in | mean | accuracy |
| Jev 1.13.0 | 0.060 (upper-case, priority) | 0.005 (0.25, priority) | 0.56 | 0.77 | 0.89 / 0.99 | 0.23 / 0.50 / 0.77 | 1.00 / 1.00 / 1.00 |
| Laya english | 0.305 (upper-case, queue) | 0.155 (0.5, queue) | 1.00 | 0.82 | 0.48 / 0.84 | 0.14 / 0.26 / 0.36 | 0.92 / 0.77 / 0.60 |
| Laya typed-dec. | 0.275 (upper-case, queue) | 0.135 (0.5, queue) | 0.76 | 0.93 | 0.43 / 0.78 | 0.26 / 0.36 / 0.45 | 0.94 / 0.85 / 0.73 |
| Kev -0.8B | 0.110 (upper-case, anger) | 0.030 (0.5, priority) | 0.96 | 0.97 | 0.55 / 0.89 | 0.39 / 0.52 / 0.65 | 0.97 / 0.99 / 0.99 |
| Kev -4B | 0.025 (5% typos, anger) | 0.015 (0.25, priority) | 0.93 | 0.94 | 0.64 / 0.90 | 0.28 / 0.50 / 0.72 | 0.99 / 0.99 / 0.99 |
| Severity ladders (F) | Anger negation (G) | Choice (G) | Threshold (G) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | monotone 3/5/10 | exact 3/5/10 | level at 0 / 4 | Spearman | acc. original | dev. expl. / para. | acc. expl. / para. | JSD | changed | loss |
| Jev 1.13.0 | 4 / 4 / 4 | 0.75 / 0.80 / 0.50 | 0.05 / 3.99 | 0.95–0.99 | 1.00 | 0.015 / 0.022 | 1.00 / 1.00 | 0.027 | 0.000 | 0.00 |
| Laya english | 3 / 1 / 1 | 0.55 / 0.45 / 0.20 | 1.22 / 2.96 | 0.81–0.91 | 0.84 | 0.604 / 0.756 | 0.19 / 0.12 | 0.012 | 0.050 | 0.06 |
| Laya typed-dec. | 2 / 2 / 2 | 0.65 / 0.45 / 0.35 | 1.33 / 3.18 | 0.89–0.96 | 0.90 | 0.382 / 0.478 | 0.28 / 0.10 | 0.014 | 0.084 | 0.23 |
| Kev -0.8B | 1 / 2 / 2 | 0.55 / 0.50 / 0.40 | 1.43 / 2.95 | 0.94–0.97 | 0.99 | 0.097 / 0.191 | 0.91 / 0.52 | 0.014 | 0.016 | 0.24 |
| Kev -4B | 2 / 3 / 3 | 0.90 / 0.75 / 0.55 | 0.52 / 3.51 | 0.96–0.98 | 0.99 | 0.036 / 0.031 | 0.99 / 0.98 | 0.001 | 0.002 | 0.11 |
| Manifest | Encoder | alone | AUROC | answers escalated | accuracy | random | requests needing a hosted call | |
|---|---|---|---|---|---|---|---|---|
| phishing | Laya english | 0.706 | 0.68 | 0.5 | 0.09 | 0.763 | 0.732 | 0.66 |
| 0.7 | 0.34 | 0.858 | 0.804 | 1.00 | ||||
| 0.9 | 0.72 | 0.954 | 0.916 | 1.00 | ||||
| phishing | Laya typed-dec. | 0.762 | 0.74 | 0.5 | 0.15 | 0.823 | 0.798 | 0.87 |
| 0.7 | 0.55 | 0.965 | 0.891 | 1.00 | ||||
| 0.9 | 0.96 | 0.998 | 0.988 | 1.00 |
| Question | Model | argmax | Bayes | value | 95% interval |
| phish.is_phishing | cheapest constant action: 4,860 | ||||
| Jev 1.13.0 | 3,200 | 4,820 | 1,620 | [ 3,740, 820] | |
| Laya english | 49,240 | 3,820 | 45,420 | [ 37,800, 53,140] | |
| Laya typed-dec. | 85,600 | 4,860 | 80,740 | [ 71,780, 90,000] | |
| Kev -0.8B | 58,000 | 4,860 | 53,140 | [ 44,960, 61,380] | |
| Kev -4B | 40,920 | 4,860 | 36,060 | [ 28,940, 43,360] | |