cs.AIOct 8, 2026

TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models

Authors: Rahul Sharma, Andrew B. Ducan, Gaétan Marceau Caron, Sebastian J. Vollmer

Organizations: DFKI GmbH · RPTU Kaiserslautern · Imperial College London · Mila - Quebec AI Institute

Abstract

System One models output calibrated probabilities over typed answers such as categorical choices, ordinal levels, or binary outcomes, via a non-generative interface. Software can act on these probabilities through thresholds, cost-weighted choices, and escalation rules. Consequently, if these probabilities are miscalibrated or wording-sensitive, the software ma take unintended actions leaving human operators with no textual rationale to inspect. Current evaluations largely report accuracy and calibration on public classification datasets without a clear reference. We present TypedBench, a benchmark for typed decision models like Jev built from seven policy-labelled generators and nine evaluation suites. We report accuracy as median and range across paraphrases, and calibration error relative to the finite-sample noise floor of a matched, perfectly calibrated predictor. We assess probability quality through selective prediction, ordinal proper scoring rules, and realised cost under asymmetric cost matrices. We evaluate a hosted model, an open encoder, and a family of open decoders spanning 0.8B-9B parameters on identical items. The hosted model follows the stated policy but is wording-sensitive and systematically underconfident; under asymmetric costs, using its probabilities can be worse than taking its top answer. Decoders route exactly and are slow as options or questions are added. The decoder is least accurate on policy questions and degrades with more options. Overall, typed decision models must be evaluated jointly on policy adherence, wording robustness, probability quality, and induced decision outcomes.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 29, 2026cs.LG

Benchmarking System One decision models against trained classifiers and language models for automated decision gates

Software that hands branching decisions to a model needs a declared option and a probability it can threshold. Typed decision models, also called System One models, return such probabilities without generating text, while supervised classifiers and generative language models are the established alternatives. One harness sends eight decision-model checkpoints from six families, including the hosted model Jev, and four open generative models from three developers the same semantic requests, and scores trained and zero-shot classifiers on the same workflow, intent, emotion and social-science items. With task labels, a fine-tuned DeBERTa-v3-large has the highest observed accuracy on every labeled benchmark but one. Without labels, no decision model is significantly more accurate than Jev on workflows or intents, but Gemma-4-31B matches it on workflows and exceeds it on CLINC-150 at higher cost and latency. Stated probabilities of generative models become unreadable when replies miss the key format, whereas key likelihoods avoid this but can saturate. A guaranteed 5 percent risk leaves Jev 0.528 of the intent decisions, and an in-scope threshold still accepts 0.310 of out-of-scope requests. On typed-decisions, swapping yes and no flips 50.5 answers per hundred for Jev and at least 16.8 for every generative model tested, against at most 6.5 for four fine-tuned decision checkpoints. Exposure to a benchmark's training data explains the largest lead of an open checkpoint, which vanishes on rater-labeled emotions. An intent-trained first stage escalating to Gemma-4-31B reaches that model's accuracy at about Jev's price. The results yield condition-dependent design rules for automated decision gates.
Sep 22, 2026cs.AI

Type-Safe Is Not Error-Free: Typed Decision Models Follow the Option Name, Not the Definition Bound to It

Typed decision models return structured results, but output-type correctness alone does not ensure that decisions follow explicit option definitions. Each option pairs a name with a definition that defines its intended meaning; the name, however, can provide a competing semantic cue. We study this conflict in Jev and two open-weight models by changing only the name-definition mapping, leaving the question, state, and the names and definition texts themselves unchanged. We measure decision flips at the level of the selected definition, rather than the returned name. On 1200 decision tasks with task-specific definitions, decision-flip rates are up to 70.4 pp higher with yes/no names than with the 0/1 control. This gap holds across all 4 binary decision rules. With yes/no names, reassignment also lowers their mean AUC from 93.8% to a below-chance 23.2%. In the binary evaluations, random strings used as option names yield mean flip rates close to those of neutral controls across all three models, with comparable balanced accuracy before reassignment. Together, these results support option-name polarity as a contributor to decision instability beyond reassignment alone. The type-error rate remains 0% throughout, showing that type-correct outputs can still fail to follow explicit option definitions.
Sep 28, 2026cs.AI

Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration

The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition AA for which the exact probability P(A)P(A) is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in \texttt{Choice} answers, P(A)P(A) and P(¬A)P(\neg A) are presented as if P(A)+P(¬A)=1P(A)+P(\neg A)=1, while a term P(U)≠0P(U)\neq0 is missing in the sum. Recovering P(U)P(U) leads to an improvement of median soft accuracy in \texttt{Choice} answers from 0.7710.771 to 0.9780.978, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.