Rejection policies must remain useful as candidate sets and tasks change. We compare Laya, Jev and Qwen2.5-7B-Instruct using public reference labels, testing Laya/Jev policy transfer at equal calibration budgets and all three models on artificial omission, natural retrieval misses and public out-of-scope queries. Source calibration often fails to preserve the target operating point. A Jev policy calibrated on DBpedia rejects 69.3% of covered Emotion test inputs, while an Emotion policy loses detection entirely. Retrieval exposes a different tradeoff: with ten intent candidates, Laya detects 99.0% of out-of-scope queries but rejects 48.8% of covered queries. Separating missing-answer sources reveals these costs alongside retrieval coverage. The benchmark provides shared inputs, explicit decision and failure categories, and reproducible scoring to assess rejection policies under the conditions in which they are reused. Code and benchmark artifacts are available at https://github.com/luckykevvv/Decision_Model_Benchmark.
Figures & tables
Task
Classes
Cal. texts
Test texts
Ordinary k
AG News
4
80
120
2, 3
DBpedia
14
280
420
2, 4, 13
Emotion
6
120
166
2, 4, 5
TREC*
6
120
150
2, 4, 5
CLINC150
150
350
400
2, 5, 10
Table 1: Sample support from official training/test sources. Classification sampling targets 20 calibration and 30 test texts per class; Emotion surprise has 16 test texts. CLINC includes 300 in-scope queries per split plus 50/100 calibration/test OOS queries. k excludes none . *TREC lacks ABBR test examples and is supplemental.
Figure 1: Evaluation protocol. Matched candidate sets define covered/missing conditions. Models receive shared inputs, with decisions recorded as correct, wrong, rejected or failed. Source calibration precedes target evaluation; target recalibration uses an equally sized target label set. Counts refer to the classification index.
Figure 2: Policy transfer at k=2 , natural names and 80 calibration pairs per task. Rows are sources and columns targets. Off-diagonal cells use source thresholds; dashed diagonals use target labels for recalibration. Target test-text n=120/420/166/150 for AG News/DBpedia/Emotion/TREC. Both metrics span 0–100%. Rates hold thresholds fixed; source data retain marginal text-bootstrap and refitted-source intervals. *TREC is supplemental.
Track / k
Model
Coverage
ncov/nmiss
Dmiss
Fcov
DOOS
Artificial / 5
Laya
–
300/300
86.0
27.7
–
Artificial / 5
Jev
–
300/300
63.7
3.3
–
Artificial / 5
Qwen7B
–
300/300
45.0
3.7
–
Artificial / 5
BM25 scope gate*
–
300/300
82.3
17.7
–
Retrieved / 2
Laya
89.7
269/31
90.3
15.2
83.0
Retrieved / 2
Jev
89.7
269/31
83.9
3.3
84.0
Table 2: Missing-answer sources in CLINC150: native explicit-NONE decisions (%). Artificial omission uses 300 matched in-scope queries; retrieval uses 300 in-scope and 100 OOS queries. Coverage is the reference-label retrieval hit rate. Dmiss/Fcov use the stated missing/covered counts; DOOS uses 100 queries. Source data retain intervals and failures. *BM25 uses 14,700 labeled utterances and fixed intent identifiers.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Frozen model/interface
Recorded configuration
Laya
Laya 0.3.21, pinned revision
CUDA, 1 CPU thread; 512 sequence / 256 head tokens; batch 8. Native temperature handling.
Jev
Jev 1.13.0, hosted Choice
Four API workers; up to 64 questions per call; native choices and rounded two-decimal scores.
Qwen7B
Qwen2.5-7B-Instruct, pinned revision [ 18 ]
BF16, greedy; batch 32; input cap 2,048 / output cap 32 tokens; exact-key parser; discrete track.
Appendix
Table 3: Recorded confirmation interfaces and configurations. Qwen7B denotes Qwen2.5-7B-Instruct. Model/checkpoint revisions and software versions are retained in runtime manifests; hosted internal compute is not inferred.
Figure 3: Native outcomes and local rejection policies on fresh-text confirmation. Natural names are used at the largest incomplete count: k=3/13/5/5 for AG News/DBpedia/Emotion/TREC. Independent test-text n=120/420/166/150 . a, covered correct, wrong, rejected and failed outcomes. b, missing-reference rejection, acceptance and failure. c, native detection (blue) and false rejection (gray), with marginal 95% text-bootstrap intervals (2,000 draws); diamonds show none-score policies selected using the full same-task calibration set. Qwen has native points only. TREC is supplemental. Native and score-threshold decisions use their respective recorded interfaces.
Task
Model
Test n
Correct
Wrong
F
D
Fail P/A
AG News
Laya
120
90.8
2.5
6.7
94.2
0.0/0.0
AG News
Jev
120
92.5
4.2
3.3
74.2
0.0/0.0
AG News
Qwen7B
120
93.3
4.2
2.5
69.2
0.0/0.0
DBpedia
Laya
420
67.1
3.8
29.0
84.8
0.0/0.0
DBpedia
Jev
420
97.4
0.7
1.9
82.4
0.0/0.0
DBpedia
Qwen7B
420
96.9
2.9
0.2
25.0
0.0/5.2
Appendix
Table 4: Reference-label confirmation with natural names at the largest incomplete count. Rates are percentages; covered correct/wrong/rejected/failure and missing rejection/acceptance/failure partitions retain all attempted test texts. Fail P/A denotes covered/missing failure percentages. *TREC has no ABBR test support and is supplemental. Marginal 95% text-bootstrap intervals are retained in source data.
Figure 4: Coverage-score reliability varies by task. Raw p(none) predicts reference-label absence on a balanced mixture with one covered and one missing response per test text. Tasks, counts and text n match Figure 3 . Ten fixed equal-width bins retain their response counts in source data; empty bins are omitted. Lines are descriptive, and the diagonal denotes agreement between mean predicted probability and observed absence fraction. Only Laya and Jev provide option probabilities. Brier uncertainty uses marginal text bootstrap; paired responses are not independent samples.
Task
Model
Brier
AUROC
Local calibrated D/F (%)
AG News
Laya
0.057
0.969
87.5/3.3
AG News
Jev
0.121
0.943
74.2/3.3
DBpedia
Laya
0.191
0.865
34.0/2.4
DBpedia
Jev
0.068
0.987
98.8/3.6
Emotion
Laya
0.326
0.636
19.3/7.8
Emotion
Jev
0.277
0.704
25.3/7.8
Appendix
Table 5: Coverage-confidence diagnostics on balanced present/absent pairs, using natural names and the largest incomplete count. Brier scores use raw p(none) ; AUROC measures ranking. Laya and Jev provide probability outputs.
Task
Model
Control
Text n
ΔD [95%]
ΔF [95%]
AG News
Laya
ID minus natural
20
-2.5 [-7.5, 0.0]
12.5 [2.5, 25.0]
AG News
Laya
Order 1 minus 0
20
-5.0 [-15.0, 0.0]
-5.0 [-15.0, 0.0]
AG News
Laya
Aligned minus old
20
0.0 [0.0, 0.0]
0.0 [0.0, 0.0]
DBpedia
Laya
ID minus natural
20
7.5 [-2.5, 17.5]
12.5 [0.0, 27.5]
DBpedia
Laya
Order 1 minus 0
20
0.0 [-15.0, 15.0]
-10.0 [-25.0, 0.0]
DBpedia
Laya
Aligned minus old
20
15.0 [0.0, 30.0]
5.0 [0.0, 15.0]
Appendix
Table 6: Matched interface controls at the largest incomplete count, using natural names except for the name contrast. Changes in D/F are percentage points with marginal paired-text 95% bootstrap intervals (2,000 draws). Each task starts with 20 test texts; both-side validity determines eligible n . Excluded comparisons remain in source data, while native failure denominators retain all attempts. Order holds members fixed; wording holds keys/order fixed. Variants are not independent texts.
Model
Source
Target
Frozen source D/F
Local reference D/F
Laya
AG News
DBpedia
89.3/5.2
90.0/6.0
Laya
AG News
Emotion
69.3/34.3
20.5/8.4
Laya
AG News
TREC*
98.7/66.7
26.7/4.0
Laya
DBpedia
AG News
85.8/4.2
85.8/3.3
Laya
DBpedia
Emotion
71.1/34.9
20.5/8.4
Laya
DBpedia
TREC*
99.3/70.0
26.7/4.0
Appendix
Table 7: All none-score task transfers at k=2 and natural names, with exactly 80 calibration pairs per source/local reference. Cells show target D/F percentages; local references consume target calibration labels. Every ordered task pair is retained. Frozen-threshold and refitted-source uncertainty are recorded separately in source data. *TREC comparisons are supplemental.
Task ( n ; k )
Model
Full acc.
D
F
D−F
AUC none
AG News (80; 3)
Laya
95.0
87.5
5.0
82.5
0.980
AG News (80; 3)
Jev
95.0
77.5
7.5
70.0
0.931
DBpedia (280; 13)
Laya
86.4
77.1
18.6
58.6
0.868
DBpedia (280; 13)
Jev
99.3
80.7
1.4
79.3
0.992
Emotion (120; 5)
Laya
51.7
68.3
45.8
22.5
0.676
Emotion (120; 5)
Jev
51.7
58.3
32.5
25.8
0.698
Appendix
Table 8: Natural-name results at the largest incomplete ordinary candidate count for each task. Full accuracy uses all classes without none. Rates and D−F are percentages/percentage points; none-score AUC compares matched present/absent requests with none. Full uncertainty intervals and all configurations appear in Appendix G .
Task
Common n
Laya D
Laya F
Jev D
Jev F
AG News
72
91.7
1.4
80.6
4.2
DBpedia
242
82.6
14.0
79.8
1.7
Emotion
47
91.5
42.6
85.1
21.3
TREC
91
100.0
69.2
24.2
0.0
Appendix
Table 9: The same texts correctly classified by both models with full natural-name options. Rates are percentages; k matches Table 8 . This selected subset supplements overall results rather than estimating population performance.
Task
Model
Cal. n
None threshold
Test D
Test F
AG News
Laya
40
0.4756
86.3
5.0
AG News
Jev
40
0.8800
70.0
1.3
DBpedia
Laya
140
0.9999
23.9
2.1
DBpedia
Jev
140
0.0600
98.9
3.2
Emotion
Laya
60
0.9846
18.3
5.8
Emotion
Jev
60
0.9800
12.5
4.2
Appendix
Table 10: Thresholds maximize calibration detection subject to empirical calibration false rejection at most 5%, with frozen test evaluation. Natural names and k match Table 8 ; reject iff p(none)>t . Test rates are percentages, not a guaranteed risk bound.
Figure 5: Discovery none-score detection–false-rejection curves on matched test pairs with natural names and k from Table 8 . Colors identify models; circles show native none choices and diamonds calibration-only selected thresholds. Native choices need not lie on a scalar none-threshold curve. Lines are empirical test diagnostics, with no test-selected recommended threshold or uncertainty band; Table 8 and the appendix provide uncertainty.
Model
Source
Target
D
F
Local D
Local F
Laya
DBpedia
AG News
81.2
2.5
91.2
6.2
Laya
Emotion
AG News
63.7
2.5
91.2
6.2
Laya
TREC
AG News
62.5
1.2
91.2
6.2
Laya
AG News
DBpedia
93.9
11.8
84.3
2.1
Laya
Emotion
DBpedia
56.8
0.7
84.3
2.1
Laya
TREC
DBpedia
54.3
0.7
84.3
2.1
Appendix
Table 11: Post-hoc cross-task transfer at k=2 with natural names and the none score. Thresholds use source calibration only. Local columns use target calibration labels. All ordered task transfers are reported; rates are percentages. These are discovery results, with marginal Wilson intervals in the accompanying JSON.
Task
Names
Model
Without none [95%]
With none [95%]
AG News
natural
Laya
95.0 [87.8, 98.0]
90.0 [81.5, 94.8]
AG News
natural
Jev
95.0 [87.8, 98.0]
90.0 [81.5, 94.8]
AG News
neutral
Laya
95.0 [87.8, 98.0]
90.0 [81.5, 94.8]
AG News
neutral
Jev
96.3 [89.5, 98.7]
90.0 [81.5, 94.8]
DBpedia
natural
Laya
86.4 [81.9, 90.0]
77.1 [71.9, 81.7]
DBpedia
natural
Jev
99.3 [97.4, 99.8]
97.5 [94.9, 98.8]
Appendix
Table 12: Full-candidate test accuracy with Wilson 95% intervals, percentages.
Group ( k , names)
Model
D [95%]
F [95%]
D−F [95%]
AG News 2 Nat.
Laya
88.8 [80.0, 94.0]
6.3 [2.7, 13.8]
82.5 [73.8, 91.3]
AG News 2 Nat.
Jev
83.8 [74.2, 90.3]
8.8 [4.3, 17.0]
75.0 [65.0, 83.8]
AG News 2 ID
Laya
91.3 [83.0, 95.7]
5.0 [2.0, 12.2]
86.3 [77.5, 93.8]
AG News 2 ID
Jev
85.0 [75.6, 91.2]
8.8 [4.3, 17.0]
76.3 [66.3, 85.0]
AG News 3 Nat.
Laya
87.5 [78.5, 93.1]
5.0 [2.0, 12.2]
82.5 [73.8, 91.3]
AG News 3 Nat.
Jev
77.5 [67.2, 85.3]
7.5 [3.5, 15.4]
70.0 [60.0, 80.0]
Appendix
Table 13: All paired native-choice rates. D/F include Wilson 95% intervals; D-F uses 2,000 paired-text bootstrap draws. Rates are percentages and differences percentage points. Nat.=natural names; ID=neutral keys.
Group ( k , names)
Model
−maxp [95%]
Entropy [95%]
None [95%]
AG News 2 Nat.
Laya
0.955 [0.911, 0.984]
0.955 [0.911, 0.984]
0.980 [0.958, 0.995]
AG News 2 Nat.
Jev
0.804 [0.721, 0.873]
0.804 [0.721, 0.873]
0.936 [0.897, 0.971]
AG News 2 ID
Laya
0.954 [0.912, 0.981]
0.954 [0.912, 0.981]
0.970 [0.941, 0.992]
AG News 2 ID
Jev
0.864 [0.793, 0.922]
0.864 [0.793, 0.922]
0.948 [0.913, 0.978]
AG News 3 Nat.
Laya
0.927 [0.872, 0.967]
0.938 [0.891, 0.973]
0.980 [0.961, 0.995]
AG News 3 Nat.
Jev
0.777 [0.696, 0.848]
0.781 [0.701, 0.849]
0.931 [0.889, 0.965]
Appendix
Table 14: All score AUROCs with 2,000 paired-text bootstrap 95% intervals. Max/entropy use plain requests; none uses requests with an added rejection option. Intervals are marginal.
Group
Model
Score
Test D
Test F
AG News 2 Nat.
Laya
−maxp
86.3
8.8
AG News 2 Nat.
Laya
Entropy
86.3
8.8
AG News 2 Nat.
Laya
None
91.3
6.3
AG News 2 Nat.
Jev
−maxp
41.3
5.0
AG News 2 Nat.
Jev
Entropy
41.3
5.0
AG News 2 Nat.
Jev
None
78.8
3.8
Appendix
Table 15: All calibration-only selected test operating points (percentages). Counts and Wilson intervals remain in frozen metrics; these are actual test rates, not a nominal 5% guarantee.
Typed decision models such as Jev offer an efficient alternative to generative LLMs in decision-making workflows by selecting directly from predefined options. When candidate sets contain no valid answer, TypeSafe recommends including an "other" or "none-of-the-above" option to enable rejection. In this report, however, we identify an arithmetic-dependent rejection bottleneck: Jev reliably selects correct numerical answers when available but frequently accepts incorrect alternatives when they are absent despite an explicit rejection option. On paired arithmetic problems, answer-present accuracy reaches 99%, while correct rejection falls to 7%. Moreover, this gap persists across numerical magnitudes, operation depths, contextual formulations, and rejection labels, and extends to scenarios such as time calculation and capacity rounding. Yet native Boolean verification achieves 99% exact-match accuracy on the same answer-absent arithmetic cases, showing that categorical rejection can fail even when the model successfully verifies candidate correctness. Finally, we show that a simple decision threshold selected on separate development problems raises arithmetic rejection accuracy from 7% to 79% while retaining 97% answer-present accuracy, substantially mitigating the failure without retraining or additional inference.
Jike Zhong, Ming Li, Yuxiang Lai
University of Southern California · University of Florida · Emory University
Software that hands branching decisions to a model needs a declared option and a probability it can threshold. Typed decision models, also called System One models, return such probabilities without generating text, while supervised classifiers and generative language models are the established alternatives. One harness sends eight decision-model checkpoints from six families, including the hosted model Jev, and four open generative models from three developers the same semantic requests, and scores trained and zero-shot classifiers on the same workflow, intent, emotion and social-science items. With task labels, a fine-tuned DeBERTa-v3-large has the highest observed accuracy on every labeled benchmark but one. Without labels, no decision model is significantly more accurate than Jev on workflows or intents, but Gemma-4-31B matches it on workflows and exceeds it on CLINC-150 at higher cost and latency. Stated probabilities of generative models become unreadable when replies miss the key format, whereas key likelihoods avoid this but can saturate. A guaranteed 5 percent risk leaves Jev 0.528 of the intent decisions, and an in-scope threshold still accepts 0.310 of out-of-scope requests. On typed-decisions, swapping yes and no flips 50.5 answers per hundred for Jev and at least 16.8 for every generative model tested, against at most 6.5 for four fine-tuned decision checkpoints. Exposure to a benchmark's training data explains the largest lead of an open checkpoint, which vanishes on rater-labeled emotions. An intent-trained first stage escalating to Gemma-4-31B reaches that model's accuracy at about Jev's price. The results yield condition-dependent design rules for automated decision gates.
Amir Rafe, Subasish Das
Texas State University, 601 University Drive, San Marcos, 78666, Texas, USA
A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable ones or ones resting on a false premise. The usual recipe thresholds a single confidence score, which cannot tell these apart. Across five instruction-tuned models from three families (2B to 14B), we find they are separate axes. Ordinary answer-confidence tracks whether an answer is right but is nearly blind to whether the question is answerable; a linear probe on hidden states does the reverse. The blind spot does not shrink with scale. It is worst on naturally occurring false-premise questions (CREPE). There, answer-confidence, P(IK), P(True), and even asking the model outright whether a premise is false all stay near chance, while a hidden-state probe reaches 0.69 to 0.77 AUROC: the model represents a problem it will not report. This turns out to be fixable. Instructing a model to check premises backfires, because it then disputes sound and false premises alike (57% false challenges), unable to tell them apart; routing the same instruction with the probe roughly triples challenge precision. We turn the two axes into a calibrated policy that answers only when an answerability score and a correctness score each clear a separately certifies behave differently: the unanswerable-answer rate is controllable at every scale, while the wrong-answer rate is capped by model accuracy, so the guarantee tightens as threshold policy certifies both budgets at 0.75 coverage of correct answers, against 0.31 for a single threshold; at 14B it is the only policy that certifies at all.