A typed decision is a choice among a fixed set of options, returned as a probability rather than as text. Systems that need typed decisions today use models trained for that purpose. This report describes AnyJev, which reads a typed decision from one prefill of a pretrained instruction-tuned language model. The readout restricts the next-token distribution at the answer position to the option tokens. It has two defects: the model assigns higher probability to some labels whatever the input, and to some positions in the option list. AnyJev corrects both with no gradient steps and no parameter changes: it divides out a label prior estimated from unlabelled inputs, and it averages log-probabilities over the K cyclic rotations of the option list. On two 20-option tasks the rotations lower the order-flip rate from 0.33 to 0.14 and from 0.33 to 0.18, and raise accuracy on 11 of 11 models on both. Reading every rotation requires K prefills. A stopping rule selected against the full-rotation decision on unlabelled states cuts that. Selecting the threshold on one unlabelled split and bounding its disagreement on a second, it reads 10.6 rotations of 18 at a verified 0.008 bound on two of four cells; selected and bounded on one split, as our serving run did, it reads 7.3 and serves 2.2 times as many decisions per second on vLLM. The code is open source.
Figures & tables
Figure 1: The same base model, three decision paths. (a) Asking for the answer in text costs one forward pass per generated token and returns prose the caller has to parse, with no probability attached. (b) AnyJev L0 runs one prefill per cyclic rotation. It restricts the next-token distribution to the option-label tokens and renormalises, corrects the label prior, maps each position back to its option, and averages in log space across the K rotations before renormalising. A, B and C are position labels whose option assignments change across rotations; the final distribution is indexed by option. Probabilities are illustrative, and one rotation’s token distribution is shown. L0 debiases the readout; confidence calibration is a separate L1 step. (c) L2-mono truncates the forward at a selected block, translates the hidden state into the final-layer space with an affine map, and uses the original output head for the option readout. The translator is fitted once on unlabelled inputs to the model’s full-depth hidden states; the exit depth is selected by agreement with its full-depth answers, without task labels. The L2-mono output is symbolic. All three paths keep the base-model weights fixed.
Figure 2: Rotation averaging on two 20-option tasks. Each row is one model, with separate subpanels for banking20 and 20 newsgroups under each metric. Hollow points show the raw readout; filled points show the rotation average of Eq. ( 3 ). Left: the order-flip rate, the share of items whose answer changes when the option list is reversed. Right: accuracy. Axes are shown as percentages. All 22 model–task pairs have a lower order-flip rate and higher accuracy after averaging. The label prior is switched off, so the changes are due to rotations alone.
banking20
20 newsgroups
injection
K=20
K=20
K=2
Accuracy
raw readout
0.665
0.609
0.720
+ rotations
0.737
0.668
0.733
+ rotations + label prior
0.752
0.675
0.762
Order-flip rate
Table 1: Means over the models that pass the answer-mass gate, 300 items per cell. The lower block counts models rather than items. Under a coin flip, 11 of 11 occurs with probability 4.9×10−4 (exact sign test, one-sided). The injection accuracy row is 3 of 7 after dropping one tie, at p=0.77 .
Figure 3: Differences on the JevBench public subset with 95% bootstrap intervals over 2 000 resamples, paired over items, one row per model. Left: debiased minus raw accuracy; every interval contains zero. Right: the change in calibration error from adding the temperature; four of the eight intervals exclude zero, and all four are improvements. Coloured intervals exclude zero.
Model
Answer mass
Raw
Debiased
ECE raw
ECE debiased
ECE + temp.
Qwen3-32B
1.000
0.798
0.789
0.140
0.128
0.091
Qwen3-8B
1.000
0.704
0.714
0.278
0.263
0.114
gpt-oss-20b
0.999
0.704
0.714
0.140
0.154
0.102
Granite-3.3-8B
0.999
0.690
0.700
0.280
0.270
0.301
Qwen2.5-7B
0.996
0.676
0.685
0.259
0.243
0.085
Mistral-7B
0.999
0.615
0.624
0.322
0.333
0.104
Table 2: JevBench public subset, 213 items scored per model. Answer mass is the normaliser of Eq. ( 1 ). ECE is expected calibration error over 15 equal-mass bins. Intervals for the differences are in Figure 3 and Table 7 .
Engine
Readout
Decisions/s
Speed-up
Rotations read
Accuracy
Agreement
vLLM
all 18 rotations
16.7
1.00
18.0
0.697
1.000
budget, wave 1
22.1
1.32
7.2
0.703
0.987
budget, wave 2
37.2
2.22
7.3
0.703
0.987
budget, wave 4
31.8
1.90
9.3
0.703
0.987
budget, wave 6
29.2
1.74
10.9
0.700
0.987
Transformers
all 18 rotations
7.0
1.00
18.0
0.697
1.000
Table 3: 300 decisions of an 18-option question on one H100. The wave is how many rotations a decision requests per engine call. Accuracy spans 0.697 to 0.703 across the rows, and agreement is measured against the full-rotation decision. Speed-up is against the all-rotations row of the same engine.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Model
mass
raw
+rot.
+prior
ECE raw
ECE +temp.
flip raw
flip +rot.
cov@5% raw
cov@5% +temp.
Granite-3.3-8B
1.000
0.657
0.810
0.807
0.293
0.052
0.353
0.143
0.033
0.523
Qwen3-8B
1.000
0.747
0.800
0.803
0.240
0.095
0.230
0.077
0.077
0.520
Phi-4-mini
0.999
0.707
0.767
0.783
0.185
0.065
0.277
0.130
0.017
0.480
Qwen3-32B
1.000
0.720
0.777
0.780
0.227
0.064
0.213
0.080
0.017
0.443
Qwen3-30B-A3B
1.000
0.730
0.757
0.770
0.249
0.086
0.143
0.103
0.277
0.470
Qwen3-4B
1.000
0.733
0.757
0.760
0.254
0.098
0.197
0.110
0.103
0.423
Appendix
Table 4: banking20 , K=20 , 300 items per model. All eleven rows pass the gate.
Model
mass
raw
+rot.
+prior
ECE raw
ECE +temp.
flip raw
flip +rot.
cov@5% raw
cov@5% +temp.
Qwen3-30B-A3B
1.000
0.737
0.743
0.740
0.242
0.096
0.140
0.087
0.303
0.003
Qwen3-32B
1.000
0.717
0.733
0.737
0.236
0.063
0.170
0.090
0.557
0.410
Qwen2.5-7B
0.998
0.663
0.703
0.710
0.273
0.096
0.237
0.133
0.047
0.407
Granite-3.3-8B
1.000
0.660
0.690
0.693
0.328
0.111
0.303
0.187
0.000
0.183
OLMo-2-7B
0.999
0.573
0.660
0.683
0.203
0.047
0.457
0.240
0.283
0.290
Mistral-7B
0.995
0.647
0.663
0.673
0.315
0.096
0.317
0.167
0.007
0.073
Appendix
Table 5: newsgroups , K=20 , 300 items per model. All eleven rows pass the gate.
Model
mass
raw
+rot.
+prior
ECE raw
ECE +temp.
flip raw
flip +rot.
cov@5% raw
cov@5% +temp.
Qwen3-32B
0.999
0.857
0.833
0.827
0.062
0.067
0.073
0.000
0.550
0.483
Phi-4-mini
0.994
0.707
0.790
0.807
0.133
0.106
0.427
0.000
0.017
0.010
Qwen3-4B
1.000
0.683
0.783
0.797
0.298
0.081
0.203
0.000
0.117
0.100
Qwen2.5-7B
0.999
0.737
0.720
0.790
0.188
0.049
0.070
0.000
0.367
0.380
Qwen3-30B-A3B
0.995
0.730
0.730
0.757
0.248
0.090
0.103
0.000
0.410
0.410
Mistral-7B
0.995
0.697
0.707
0.733
0.167
0.069
0.333
0.000
0.017
0.143
Appendix
Table 6: injection , K=2 , 300 items per model. Three rows fall below the answer-mass gate and are excluded from every statistic reported for this task. The rotation average removes order dependence on every row that passes.
Model
raw
debiased
diff.
95% CI
ECE deb.
ECE +temp.
diff.
95% CI
Qwen3-32B
0.798
0.789
-0.009
[-0.042, +0.019]
0.128
0.091
-0.037
[-0.082, +0.010]
Qwen3-8B
0.704
0.714
+0.009
[-0.009, +0.028]
0.263
0.114
-0.149
[-0.179, -0.076]
gpt-oss-20b
0.704
0.714
+0.009
[-0.033, +0.052]
0.154
0.102
-0.052
[-0.092, +0.004]
Granite-3.3-8B
0.690
0.700
+0.009
[-0.019, +0.038]
0.270
0.301
+0.030
[-0.038, +0.124]
Qwen2.5-7B
0.676
0.685
+0.009
[-0.009, +0.033]
0.243
0.085
-0.159
[-0.177, -0.068]
Mistral-7B
0.615
0.624
+0.009
[-0.019, +0.038]
0.333
0.104
-0.229
[-0.251, -0.121]
Appendix
Table 7: Per-model differences with 95% bootstrap intervals over 2 000 resamples, paired over items. The temperature is refitted inside every resample. Reusing a fit made on the full sample would place the whole sample inside every draw and narrow the interval.
Options K
Items
Decisions
Raw
Debiased
Difference
95% CI
2
74
592
0.677
0.676
-0.002
[-0.017, +0.015]
3
15
120
0.333
0.317
-0.017
[-0.050, +0.017]
4
53
424
0.649
0.665
+0.017
[-0.012, +0.045]
5
55
440
0.805
0.811
+0.007
[-0.016, +0.030]
6
16
128
0.461
0.492
+0.031
[+0.008, +0.062]
Slope of the difference on K
+0.0059
[-0.0012, +0.0132]
Appendix
Table 8: The difference against the number of options, all eight models pooled. Resampling is clustered by item, so a drawn item contributes all eight models’ decisions. The final two rows replace those five interval tests with a single test of the trend.
Model
easy (48)
original (72)
hard (111)
Qwen3-32B
1.000
0.967
0.590
Qwen3-8B
1.000
0.817
0.524
gpt-oss-20b
1.000
0.850
0.505
Granite-3.3-8B
1.000
0.900
0.448
Qwen2.5-7B
1.000
0.867
0.438
Mistral-7B
1.000
0.717
0.400
Appendix
Table 9: Accuracy by difficulty tier of the public subset. These are the three public files, not the partition the benchmark’s overall score uses. The counts in the header are the published items per file; the accuracies are over the scored items only, which are 48 easy, 60 original and 105 hard, because the 18 unscored items fall in the second and third files.
Figure 4: Left: every cyclic rotation scored on its own, against the full-rotation average, on four (model, task) cells. On all four the average falls at or below the best single rotation, by 0.3 to 2.3 points, and above the worst by 3.5 to 9.3 points. Right: where the budget stops on one 18-option cell, over 600 decisions, for a rule with no minimum-rotation floor; the shipped default reads two rotations before it may stop, which moves the leftmost bar.
Model
Task
K
Threshold
Verify rate
Verify bound
Clears 1%
Rotations
vs. K
Qwen2.5-7B
massive_route
18
10.25
0.0017
0.0079
yes
10.61
1.70 ×
Qwen2.5-7B
newsgroups
20
9.50
0.0017
0.0079
yes
9.01
2.22 ×
Qwen3-8B
massive_route
18
—
—
—
—
—
—
Qwen3-8B
newsgroups
20
9.00
0.0033
0.0105
no
5.31
3.77 ×
Appendix
Table 10: The certificate with the threshold selected on one third of an unlabelled batch and the bound computed on the other two thirds, for that one threshold. The verify bound carries no selection multiplicity. Two of the four cells clear a 1% target this way; on massive_route with Qwen3-8B no threshold on the grid cleared it on the selection split.
Stopping statistic
Cells certified
Mean rotations saved
log-odds margin alone
4 of 4
3.46 ×
log-odds margin + unanimity
4 of 4
2.91 ×
probability gap alone
2 of 4
3.19 ×
probability gap + unanimity
2 of 4
2.64 ×
Appendix
Table 11: Which stopping statistic satisfies Eq. ( 5 ) at ε=0.01 , over four (model, task) cells. A rule with fewer certified cells did not run worse: no threshold on the grid cleared the target on those cells at any budget.
Stopping statistic
Model
Task
K
Threshold
Rotations
vs. K
Disagreement
log-odds margin alone
Qwen2.5-7B
massive_route
18
4.5
6.12
2.94 ×
0.0083
Qwen2.5-7B
newsgroups
20
6.0
6.00
3.34 ×
0.0000
Qwen3-8B
massive_route
18
9.5
5.37
3.36 ×
0.0000
Qwen3-8B
newsgroups
20
8.25
4.75
4.21 ×
0.0033
log-odds margin + unanimity
Qwen2.5-7B
massive_route
18
4.25
7.96
2.26 ×
0.0083
Qwen2.5-7B
newsgroups
20
6.0
6.63
3.02 ×
0.0000
Appendix
Table 12: Every cell each stopping rule certifies at ε=0.01 , with its threshold and the rotations it then reads. Rules with two rows did not certify the other two cells at any threshold on the grid. These thresholds come from a different calibration sample than the serving run of Section 7 , which is why the same cell appears there at 6.0.
Model
Blocks
Depth at 0.98
Depth at 0.95
Before crossing
At crossing
Agreement at full depth
Mistral-7B
32
66%
66%
0.320 at 59%
0.985 at 66%
1.000
Qwen3-8B
36
89%
69%
0.087 at 64%
0.965 at 69%
0.993
Qwen3-1.7B
28
100%
79%
0.188 at 75%
0.958 at 79%
0.990
Qwen2.5-7B
28
100%
89%
0.420 at 79%
0.950 at 89%
0.983
Qwen3-32B
64
100%
91%
0.228 at 80%
0.973 at 91%
0.990
gpt-oss-20b
24
100%
92%
0.207 at 79%
0.950 at 92%
0.983
Appendix
Table 13: Agreement with the model’s own full-depth decision, for a readout taken at a fraction of the blocks and mapped into the final basis. Before crossing and at crossing are the last depth measured below 0.90 agreement and the first at or above it.
Figure 5: Agreement against depth, eight models. Markers are the depths we ran, and the curve is drawn through them rather than smoothed, because the intermediate depths were not measured. Hollow markers show the first depth reaching 0.95. Dashed lines mark 0.95 and 0.98.
System / readout
Overall ↑
95% CI
Choice
Yes/no
Ordinal
ECE ↓
n=18,808
n=17,572
n=9,940
(shard mean)
Reference rules (not learned systems)
Uniform random (expected)
34.29
—
25.77
50.00
22.63
—
Majority by question (test oracle)
54.28
—
41.42
73.57
44.51
—
Published baseline interfaces
Bespoke-Nimble-9B
70.57
[70.16, 70.99]
73.60
81.06
46.33
0.0889
Appendix
Table 14: All full-split BEV Decision Mix results. Accuracy and its intervals are in percent; ECE is on a 0–1 scale. Higher accuracy and lower ECE are better. Bold marks the best measured system in each metric, including ties.
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
Dedicated decision models such as Jev map unstructured language to probability distributions over finite choices, allowing their outputs to directly route requests, select tools, and trigger actions. Yet real-world inputs rarely arrive in isolation: they come with background details and surrounding context. We find that short additions that fit naturally into this context can nevertheless redirect an otherwise correct decision, even when the correct answer remains unchanged. To study this behavior, we fix a wrong target option for each initially correct item and use the model's option probabilities to refine fluent context additions while preserving the source, question, choices, and gold answer. Within 64 accepted target evaluations, the optimizer identifies contexts that redirect Jev on 312 of 508 initially correct decisions (61.4%); in 229 cases, Jev assigns at least 0.7 probability to the fixed wrong option. Across seven datasets, three additional decision systems show targeted flip rates of 64.9%-73.2% on decisions they initially answer correctly. Taken together, these results expose a pronounced fragility in current decision models: short, ordinary-looking context can shift a correct choice to a high-confidence wrong one. Because these models turn language directly into downstream choices, this sensitivity raises concerns about treating their probability outputs as reliable decision interfaces.
System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper, we introduce Chinese-Jev, a System One model that addresses this gap through a unified data processing and training pipeline. Our data processing protocol converts heterogeneous Chinese-language annotations into probability targets over candidate options, enabling a shared training formulation across domains and question formats. To enable efficient inference, Chinese-Jev adopts a lightweight encoder-only backbone for text encoding and learns to score candidate answers through decision-oriented training. To address the misalignment between the pre-training distribution and downstream Chinese-language scenarios, we first train the model on a general-purpose corpus of 10 million examples, then fine-tune it separately for the medical, legal, and financial domains. To evaluate decision accuracy and calibration in both general and domain-specific Chinese-language settings, we introduce Chinese-Jev Bench (CJ-Bench). After first-stage pre-training, Chinese-Jev exceeds the accuracy of the closed-source Jev model by 1.24% on general-domain tasks while achieving a 20.3x speedup. Subsequent domain-specific fine-tuning yields a 4.0% accuracy improvement over Jev in medicine and achieves 92% of Jev's average accuracy across specialized domains, with a 17x speedup and an average latency of only 15 ms per example. We further demonstrate on-device deployment of an INT8-quantized model on mobile devices, achieving an inference latency of approximately 1.0 second per decision. The project is available at https://gulucaptain.github.io/Chinese-Jev/.
Zexiao Wang, Zihao Zhang, Xudong Wang +5
Fudan University · National University of Singapore · University of Chinese Academy of Sciences