Rejection policies must remain useful as candidate sets and tasks change. We compare Laya, Jev and Qwen2.5-7B-Instruct using public reference labels, testing Laya/Jev policy transfer at equal calibration budgets and all three models on artificial omission, natural retrieval misses and public out-of-scope queries. Source calibration often fails to preserve the target operating point. A Jev policy calibrated on DBpedia rejects 69.3% of covered Emotion test inputs, while an Emotion policy loses detection entirely. Retrieval exposes a different tradeoff: with ten intent candidates, Laya detects 99.0% of out-of-scope queries but rejects 48.8% of covered queries. Separating missing-answer sources reveals these costs alongside retrieval coverage. The benchmark provides shared inputs, explicit decision and failure categories, and reproducible scoring to assess rejection policies under the conditions in which they are reused. Code and benchmark artifacts are available at https://github.com/luckykevvv/Decision_Model_Benchmark.
Figures & tables
Task
Classes
Cal. texts
Test texts
Ordinary k
AG News
4
80
120
2, 3
DBpedia
14
280
420
2, 4, 13
Emotion
6
120
166
2, 4, 5
TREC*
6
120
150
2, 4, 5
CLINC150
150
350
400
2, 5, 10
Table 1: Sample support from official training/test sources. Classification sampling targets 20 calibration and 30 test texts per class; Emotion surprise has 16 test texts. CLINC includes 300 in-scope queries per split plus 50/100 calibration/test OOS queries. k excludes none . *TREC lacks ABBR test examples and is supplemental.
Figure 1: Evaluation protocol. Matched candidate sets define covered/missing conditions. Models receive shared inputs, with decisions recorded as correct, wrong, rejected or failed. Source calibration precedes target evaluation; target recalibration uses an equally sized target label set. Counts refer to the classification index.
Figure 2: Policy transfer at k=2 , natural names and 80 calibration pairs per task. Rows are sources and columns targets. Off-diagonal cells use source thresholds; dashed diagonals use target labels for recalibration. Target test-text n=120/420/166/150 for AG News/DBpedia/Emotion/TREC. Both metrics span 0–100%. Rates hold thresholds fixed; source data retain marginal text-bootstrap and refitted-source intervals. *TREC is supplemental.
Track / k
Model
Coverage
ncov/nmiss
Dmiss
Fcov
DOOS
Artificial / 5
Laya
–
300/300
86.0
27.7
–
Artificial / 5
Jev
–
300/300
63.7
3.3
–
Artificial / 5
Qwen7B
–
300/300
45.0
3.7
–
Artificial / 5
BM25 scope gate*
–
300/300
82.3
17.7
–
Retrieved / 2
Laya
89.7
269/31
90.3
15.2
83.0
Retrieved / 2
Jev
89.7
269/31
83.9
3.3
84.0
Table 2: Missing-answer sources in CLINC150: native explicit-NONE decisions (%). Artificial omission uses 300 matched in-scope queries; retrieval uses 300 in-scope and 100 OOS queries. Coverage is the reference-label retrieval hit rate. Dmiss/Fcov use the stated missing/covered counts; DOOS uses 100 queries. Source data retain intervals and failures. *BM25 uses 14,700 labeled utterances and fixed intent identifiers.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Frozen model/interface
Recorded configuration
Laya
Laya 0.3.21, pinned revision
CUDA, 1 CPU thread; 512 sequence / 256 head tokens; batch 8. Native temperature handling.
Jev
Jev 1.13.0, hosted Choice
Four API workers; up to 64 questions per call; native choices and rounded two-decimal scores.
Qwen7B
Qwen2.5-7B-Instruct, pinned revision [ 18 ]
BF16, greedy; batch 32; input cap 2,048 / output cap 32 tokens; exact-key parser; discrete track.
Appendix
Table 3: Recorded confirmation interfaces and configurations. Qwen7B denotes Qwen2.5-7B-Instruct. Model/checkpoint revisions and software versions are retained in runtime manifests; hosted internal compute is not inferred.
Figure 3: Native outcomes and local rejection policies on fresh-text confirmation. Natural names are used at the largest incomplete count: k=3/13/5/5 for AG News/DBpedia/Emotion/TREC. Independent test-text n=120/420/166/150 . a, covered correct, wrong, rejected and failed outcomes. b, missing-reference rejection, acceptance and failure. c, native detection (blue) and false rejection (gray), with marginal 95% text-bootstrap intervals (2,000 draws); diamonds show none-score policies selected using the full same-task calibration set. Qwen has native points only. TREC is supplemental. Native and score-threshold decisions use their respective recorded interfaces.
Task
Model
Test n
Correct
Wrong
F
D
Fail P/A
AG News
Laya
120
90.8
2.5
6.7
94.2
0.0/0.0
AG News
Jev
120
92.5
4.2
3.3
74.2
0.0/0.0
AG News
Qwen7B
120
93.3
4.2
2.5
69.2
0.0/0.0
DBpedia
Laya
420
67.1
3.8
29.0
84.8
0.0/0.0
DBpedia
Jev
420
97.4
0.7
1.9
82.4
0.0/0.0
DBpedia
Qwen7B
420
96.9
2.9
0.2
25.0
0.0/5.2
Appendix
Table 4: Reference-label confirmation with natural names at the largest incomplete count. Rates are percentages; covered correct/wrong/rejected/failure and missing rejection/acceptance/failure partitions retain all attempted test texts. Fail P/A denotes covered/missing failure percentages. *TREC has no ABBR test support and is supplemental. Marginal 95% text-bootstrap intervals are retained in source data.
Figure 4: Coverage-score reliability varies by task. Raw p(none) predicts reference-label absence on a balanced mixture with one covered and one missing response per test text. Tasks, counts and text n match Figure 3 . Ten fixed equal-width bins retain their response counts in source data; empty bins are omitted. Lines are descriptive, and the diagonal denotes agreement between mean predicted probability and observed absence fraction. Only Laya and Jev provide option probabilities. Brier uncertainty uses marginal text bootstrap; paired responses are not independent samples.
Task
Model
Brier
AUROC
Local calibrated D/F (%)
AG News
Laya
0.057
0.969
87.5/3.3
AG News
Jev
0.121
0.943
74.2/3.3
DBpedia
Laya
0.191
0.865
34.0/2.4
DBpedia
Jev
0.068
0.987
98.8/3.6
Emotion
Laya
0.326
0.636
19.3/7.8
Emotion
Jev
0.277
0.704
25.3/7.8
Appendix
Table 5: Coverage-confidence diagnostics on balanced present/absent pairs, using natural names and the largest incomplete count. Brier scores use raw p(none) ; AUROC measures ranking. Laya and Jev provide probability outputs.
Task
Model
Control
Text n
ΔD [95%]
ΔF [95%]
AG News
Laya
ID minus natural
20
-2.5 [-7.5, 0.0]
12.5 [2.5, 25.0]
AG News
Laya
Order 1 minus 0
20
-5.0 [-15.0, 0.0]
-5.0 [-15.0, 0.0]
AG News
Laya
Aligned minus old
20
0.0 [0.0, 0.0]
0.0 [0.0, 0.0]
DBpedia
Laya
ID minus natural
20
7.5 [-2.5, 17.5]
12.5 [0.0, 27.5]
DBpedia
Laya
Order 1 minus 0
20
0.0 [-15.0, 15.0]
-10.0 [-25.0, 0.0]
DBpedia
Laya
Aligned minus old
20
15.0 [0.0, 30.0]
5.0 [0.0, 15.0]
Appendix
Table 6: Matched interface controls at the largest incomplete count, using natural names except for the name contrast. Changes in D/F are percentage points with marginal paired-text 95% bootstrap intervals (2,000 draws). Each task starts with 20 test texts; both-side validity determines eligible n . Excluded comparisons remain in source data, while native failure denominators retain all attempts. Order holds members fixed; wording holds keys/order fixed. Variants are not independent texts.
Model
Source
Target
Frozen source D/F
Local reference D/F
Laya
AG News
DBpedia
89.3/5.2
90.0/6.0
Laya
AG News
Emotion
69.3/34.3
20.5/8.4
Laya
AG News
TREC*
98.7/66.7
26.7/4.0
Laya
DBpedia
AG News
85.8/4.2
85.8/3.3
Laya
DBpedia
Emotion
71.1/34.9
20.5/8.4
Laya
DBpedia
TREC*
99.3/70.0
26.7/4.0
Appendix
Table 7: All none-score task transfers at k=2 and natural names, with exactly 80 calibration pairs per source/local reference. Cells show target D/F percentages; local references consume target calibration labels. Every ordered task pair is retained. Frozen-threshold and refitted-source uncertainty are recorded separately in source data. *TREC comparisons are supplemental.
Task ( n ; k )
Model
Full acc.
D
F
D−F
AUC none
AG News (80; 3)
Laya
95.0
87.5
5.0
82.5
0.980
AG News (80; 3)
Jev
95.0
77.5
7.5
70.0
0.931
DBpedia (280; 13)
Laya
86.4
77.1
18.6
58.6
0.868
DBpedia (280; 13)
Jev
99.3
80.7
1.4
79.3
0.992
Emotion (120; 5)
Laya
51.7
68.3
45.8
22.5
0.676
Emotion (120; 5)
Jev
51.7
58.3
32.5
25.8
0.698
Appendix
Table 8: Natural-name results at the largest incomplete ordinary candidate count for each task. Full accuracy uses all classes without none. Rates and D−F are percentages/percentage points; none-score AUC compares matched present/absent requests with none. Full uncertainty intervals and all configurations appear in Appendix G .
Task
Common n
Laya D
Laya F
Jev D
Jev F
AG News
72
91.7
1.4
80.6
4.2
DBpedia
242
82.6
14.0
79.8
1.7
Emotion
47
91.5
42.6
85.1
21.3
TREC
91
100.0
69.2
24.2
0.0
Appendix
Table 9: The same texts correctly classified by both models with full natural-name options. Rates are percentages; k matches Table 8 . This selected subset supplements overall results rather than estimating population performance.
Task
Model
Cal. n
None threshold
Test D
Test F
AG News
Laya
40
0.4756
86.3
5.0
AG News
Jev
40
0.8800
70.0
1.3
DBpedia
Laya
140
0.9999
23.9
2.1
DBpedia
Jev
140
0.0600
98.9
3.2
Emotion
Laya
60
0.9846
18.3
5.8
Emotion
Jev
60
0.9800
12.5
4.2
Appendix
Table 10: Thresholds maximize calibration detection subject to empirical calibration false rejection at most 5%, with frozen test evaluation. Natural names and k match Table 8 ; reject iff p(none)>t . Test rates are percentages, not a guaranteed risk bound.
Figure 5: Discovery none-score detection–false-rejection curves on matched test pairs with natural names and k from Table 8 . Colors identify models; circles show native none choices and diamonds calibration-only selected thresholds. Native choices need not lie on a scalar none-threshold curve. Lines are empirical test diagnostics, with no test-selected recommended threshold or uncertainty band; Table 8 and the appendix provide uncertainty.
Model
Source
Target
D
F
Local D
Local F
Laya
DBpedia
AG News
81.2
2.5
91.2
6.2
Laya
Emotion
AG News
63.7
2.5
91.2
6.2
Laya
TREC
AG News
62.5
1.2
91.2
6.2
Laya
AG News
DBpedia
93.9
11.8
84.3
2.1
Laya
Emotion
DBpedia
56.8
0.7
84.3
2.1
Laya
TREC
DBpedia
54.3
0.7
84.3
2.1
Appendix
Table 11: Post-hoc cross-task transfer at k=2 with natural names and the none score. Thresholds use source calibration only. Local columns use target calibration labels. All ordered task transfers are reported; rates are percentages. These are discovery results, with marginal Wilson intervals in the accompanying JSON.
Task
Names
Model
Without none [95%]
With none [95%]
AG News
natural
Laya
95.0 [87.8, 98.0]
90.0 [81.5, 94.8]
AG News
natural
Jev
95.0 [87.8, 98.0]
90.0 [81.5, 94.8]
AG News
neutral
Laya
95.0 [87.8, 98.0]
90.0 [81.5, 94.8]
AG News
neutral
Jev
96.3 [89.5, 98.7]
90.0 [81.5, 94.8]
DBpedia
natural
Laya
86.4 [81.9, 90.0]
77.1 [71.9, 81.7]
DBpedia
natural
Jev
99.3 [97.4, 99.8]
97.5 [94.9, 98.8]
Appendix
Table 12: Full-candidate test accuracy with Wilson 95% intervals, percentages.
Group ( k , names)
Model
D [95%]
F [95%]
D−F [95%]
AG News 2 Nat.
Laya
88.8 [80.0, 94.0]
6.3 [2.7, 13.8]
82.5 [73.8, 91.3]
AG News 2 Nat.
Jev
83.8 [74.2, 90.3]
8.8 [4.3, 17.0]
75.0 [65.0, 83.8]
AG News 2 ID
Laya
91.3 [83.0, 95.7]
5.0 [2.0, 12.2]
86.3 [77.5, 93.8]
AG News 2 ID
Jev
85.0 [75.6, 91.2]
8.8 [4.3, 17.0]
76.3 [66.3, 85.0]
AG News 3 Nat.
Laya
87.5 [78.5, 93.1]
5.0 [2.0, 12.2]
82.5 [73.8, 91.3]
AG News 3 Nat.
Jev
77.5 [67.2, 85.3]
7.5 [3.5, 15.4]
70.0 [60.0, 80.0]
Appendix
Table 13: All paired native-choice rates. D/F include Wilson 95% intervals; D-F uses 2,000 paired-text bootstrap draws. Rates are percentages and differences percentage points. Nat.=natural names; ID=neutral keys.
Group ( k , names)
Model
−maxp [95%]
Entropy [95%]
None [95%]
AG News 2 Nat.
Laya
0.955 [0.911, 0.984]
0.955 [0.911, 0.984]
0.980 [0.958, 0.995]
AG News 2 Nat.
Jev
0.804 [0.721, 0.873]
0.804 [0.721, 0.873]
0.936 [0.897, 0.971]
AG News 2 ID
Laya
0.954 [0.912, 0.981]
0.954 [0.912, 0.981]
0.970 [0.941, 0.992]
AG News 2 ID
Jev
0.864 [0.793, 0.922]
0.864 [0.793, 0.922]
0.948 [0.913, 0.978]
AG News 3 Nat.
Laya
0.927 [0.872, 0.967]
0.938 [0.891, 0.973]
0.980 [0.961, 0.995]
AG News 3 Nat.
Jev
0.777 [0.696, 0.848]
0.781 [0.701, 0.849]
0.931 [0.889, 0.965]
Appendix
Table 14: All score AUROCs with 2,000 paired-text bootstrap 95% intervals. Max/entropy use plain requests; none uses requests with an added rejection option. Intervals are marginal.
Group
Model
Score
Test D
Test F
AG News 2 Nat.
Laya
−maxp
86.3
8.8
AG News 2 Nat.
Laya
Entropy
86.3
8.8
AG News 2 Nat.
Laya
None
91.3
6.3
AG News 2 Nat.
Jev
−maxp
41.3
5.0
AG News 2 Nat.
Jev
Entropy
41.3
5.0
AG News 2 Nat.
Jev
None
78.8
3.8
Appendix
Table 15: All calibration-only selected test operating points (percentages). Counts and Wilson intervals remain in frozen metrics; these are actual test rates, not a nominal 5% guarantee.