A typed decision is a choice among a fixed set of options, returned as a probability rather than as text. Systems that need typed decisions today use models trained for that purpose. This report describes AnyJev, which reads a typed decision from one prefill of a pretrained instruction-tuned language model. The readout restricts the next-token distribution at the answer position to the option tokens. It has two defects: the model assigns higher probability to some labels whatever the input, and to some positions in the option list. AnyJev corrects both with no gradient steps and no parameter changes: it divides out a label prior estimated from unlabelled inputs, and it averages log-probabilities over the K cyclic rotations of the option list. On two 20-option tasks the rotations lower the order-flip rate from 0.33 to 0.14 and from 0.33 to 0.18, and raise accuracy on 11 of 11 models on both. Reading every rotation requires K prefills. A stopping rule selected against the full-rotation decision on unlabelled states cuts that. Selecting the threshold on one unlabelled split and bounding its disagreement on a second, it reads 10.6 rotations of 18 at a verified 0.008 bound on two of four cells; selected and bounded on one split, as our serving run did, it reads 7.3 and serves 2.2 times as many decisions per second on vLLM. The code is open source.
Figures & tables
Figure 1: The same base model, three decision paths. (a) Asking for the answer in text costs one forward pass per generated token and returns prose the caller has to parse, with no probability attached. (b) AnyJev L0 runs one prefill per cyclic rotation. It restricts the next-token distribution to the option-label tokens and renormalises, corrects the label prior, maps each position back to its option, and averages in log space across the K rotations before renormalising. A, B and C are position labels whose option assignments change across rotations; the final distribution is indexed by option. Probabilities are illustrative, and one rotation’s token distribution is shown. L0 debiases the readout; confidence calibration is a separate L1 step. (c) L2-mono truncates the forward at a selected block, translates the hidden state into the final-layer space with an affine map, and uses the original output head for the option readout. The translator is fitted once on unlabelled inputs to the model’s full-depth hidden states; the exit depth is selected by agreement with its full-depth answers, without task labels. The L2-mono output is symbolic. All three paths keep the base-model weights fixed.
Figure 2: Rotation averaging on two 20-option tasks. Each row is one model, with separate subpanels for banking20 and 20 newsgroups under each metric. Hollow points show the raw readout; filled points show the rotation average of Eq. ( 3 ). Left: the order-flip rate, the share of items whose answer changes when the option list is reversed. Right: accuracy. Axes are shown as percentages. All 22 model–task pairs have a lower order-flip rate and higher accuracy after averaging. The label prior is switched off, so the changes are due to rotations alone.
banking20
20 newsgroups
injection
K=20
K=20
K=2
Accuracy
raw readout
0.665
0.609
0.720
+ rotations
0.737
0.668
0.733
+ rotations + label prior
0.752
0.675
0.762
Order-flip rate
Table 1: Means over the models that pass the answer-mass gate, 300 items per cell. The lower block counts models rather than items. Under a coin flip, 11 of 11 occurs with probability 4.9×10−4 (exact sign test, one-sided). The injection accuracy row is 3 of 7 after dropping one tie, at p=0.77 .
Figure 3: Differences on the JevBench public subset with 95% bootstrap intervals over 2 000 resamples, paired over items, one row per model. Left: debiased minus raw accuracy; every interval contains zero. Right: the change in calibration error from adding the temperature; four of the eight intervals exclude zero, and all four are improvements. Coloured intervals exclude zero.
Model
Answer mass
Raw
Debiased
ECE raw
ECE debiased
ECE + temp.
Qwen3-32B
1.000
0.798
0.789
0.140
0.128
0.091
Qwen3-8B
1.000
0.704
0.714
0.278
0.263
0.114
gpt-oss-20b
0.999
0.704
0.714
0.140
0.154
0.102
Granite-3.3-8B
0.999
0.690
0.700
0.280
0.270
0.301
Qwen2.5-7B
0.996
0.676
0.685
0.259
0.243
0.085
Mistral-7B
0.999
0.615
0.624
0.322
0.333
0.104
Table 2: JevBench public subset, 213 items scored per model. Answer mass is the normaliser of Eq. ( 1 ). ECE is expected calibration error over 15 equal-mass bins. Intervals for the differences are in Figure 3 and Table 7 .
Engine
Readout
Decisions/s
Speed-up
Rotations read
Accuracy
Agreement
vLLM
all 18 rotations
16.7
1.00
18.0
0.697
1.000
budget, wave 1
22.1
1.32
7.2
0.703
0.987
budget, wave 2
37.2
2.22
7.3
0.703
0.987
budget, wave 4
31.8
1.90
9.3
0.703
0.987
budget, wave 6
29.2
1.74
10.9
0.700
0.987
Transformers
all 18 rotations
7.0
1.00
18.0
0.697
1.000
Table 3: 300 decisions of an 18-option question on one H100. The wave is how many rotations a decision requests per engine call. Accuracy spans 0.697 to 0.703 across the rows, and agreement is measured against the full-rotation decision. Speed-up is against the all-rotations row of the same engine.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Model
mass
raw
+rot.
+prior
ECE raw
ECE +temp.
flip raw
flip +rot.
cov@5% raw
cov@5% +temp.
Granite-3.3-8B
1.000
0.657
0.810
0.807
0.293
0.052
0.353
0.143
0.033
0.523
Qwen3-8B
1.000
0.747
0.800
0.803
0.240
0.095
0.230
0.077
0.077
0.520
Phi-4-mini
0.999
0.707
0.767
0.783
0.185
0.065
0.277
0.130
0.017
0.480
Qwen3-32B
1.000
0.720
0.777
0.780
0.227
0.064
0.213
0.080
0.017
0.443
Qwen3-30B-A3B
1.000
0.730
0.757
0.770
0.249
0.086
0.143
0.103
0.277
0.470
Qwen3-4B
1.000
0.733
0.757
0.760
0.254
0.098
0.197
0.110
0.103
0.423
Appendix
Table 4: banking20 , K=20 , 300 items per model. All eleven rows pass the gate.
Model
mass
raw
+rot.
+prior
ECE raw
ECE +temp.
flip raw
flip +rot.
cov@5% raw
cov@5% +temp.
Qwen3-30B-A3B
1.000
0.737
0.743
0.740
0.242
0.096
0.140
0.087
0.303
0.003
Qwen3-32B
1.000
0.717
0.733
0.737
0.236
0.063
0.170
0.090
0.557
0.410
Qwen2.5-7B
0.998
0.663
0.703
0.710
0.273
0.096
0.237
0.133
0.047
0.407
Granite-3.3-8B
1.000
0.660
0.690
0.693
0.328
0.111
0.303
0.187
0.000
0.183
OLMo-2-7B
0.999
0.573
0.660
0.683
0.203
0.047
0.457
0.240
0.283
0.290
Mistral-7B
0.995
0.647
0.663
0.673
0.315
0.096
0.317
0.167
0.007
0.073
Appendix
Table 5: newsgroups , K=20 , 300 items per model. All eleven rows pass the gate.
Model
mass
raw
+rot.
+prior
ECE raw
ECE +temp.
flip raw
flip +rot.
cov@5% raw
cov@5% +temp.
Qwen3-32B
0.999
0.857
0.833
0.827
0.062
0.067
0.073
0.000
0.550
0.483
Phi-4-mini
0.994
0.707
0.790
0.807
0.133
0.106
0.427
0.000
0.017
0.010
Qwen3-4B
1.000
0.683
0.783
0.797
0.298
0.081
0.203
0.000
0.117
0.100
Qwen2.5-7B
0.999
0.737
0.720
0.790
0.188
0.049
0.070
0.000
0.367
0.380
Qwen3-30B-A3B
0.995
0.730
0.730
0.757
0.248
0.090
0.103
0.000
0.410
0.410
Mistral-7B
0.995
0.697
0.707
0.733
0.167
0.069
0.333
0.000
0.017
0.143
Appendix
Table 6: injection , K=2 , 300 items per model. Three rows fall below the answer-mass gate and are excluded from every statistic reported for this task. The rotation average removes order dependence on every row that passes.
Model
raw
debiased
diff.
95% CI
ECE deb.
ECE +temp.
diff.
95% CI
Qwen3-32B
0.798
0.789
-0.009
[-0.042, +0.019]
0.128
0.091
-0.037
[-0.082, +0.010]
Qwen3-8B
0.704
0.714
+0.009
[-0.009, +0.028]
0.263
0.114
-0.149
[-0.179, -0.076]
gpt-oss-20b
0.704
0.714
+0.009
[-0.033, +0.052]
0.154
0.102
-0.052
[-0.092, +0.004]
Granite-3.3-8B
0.690
0.700
+0.009
[-0.019, +0.038]
0.270
0.301
+0.030
[-0.038, +0.124]
Qwen2.5-7B
0.676
0.685
+0.009
[-0.009, +0.033]
0.243
0.085
-0.159
[-0.177, -0.068]
Mistral-7B
0.615
0.624
+0.009
[-0.019, +0.038]
0.333
0.104
-0.229
[-0.251, -0.121]
Appendix
Table 7: Per-model differences with 95% bootstrap intervals over 2 000 resamples, paired over items. The temperature is refitted inside every resample. Reusing a fit made on the full sample would place the whole sample inside every draw and narrow the interval.
Options K
Items
Decisions
Raw
Debiased
Difference
95% CI
2
74
592
0.677
0.676
-0.002
[-0.017, +0.015]
3
15
120
0.333
0.317
-0.017
[-0.050, +0.017]
4
53
424
0.649
0.665
+0.017
[-0.012, +0.045]
5
55
440
0.805
0.811
+0.007
[-0.016, +0.030]
6
16
128
0.461
0.492
+0.031
[+0.008, +0.062]
Slope of the difference on K
+0.0059
[-0.0012, +0.0132]
Appendix
Table 8: The difference against the number of options, all eight models pooled. Resampling is clustered by item, so a drawn item contributes all eight models’ decisions. The final two rows replace those five interval tests with a single test of the trend.
Model
easy (48)
original (72)
hard (111)
Qwen3-32B
1.000
0.967
0.590
Qwen3-8B
1.000
0.817
0.524
gpt-oss-20b
1.000
0.850
0.505
Granite-3.3-8B
1.000
0.900
0.448
Qwen2.5-7B
1.000
0.867
0.438
Mistral-7B
1.000
0.717
0.400
Appendix
Table 9: Accuracy by difficulty tier of the public subset. These are the three public files, not the partition the benchmark’s overall score uses. The counts in the header are the published items per file; the accuracies are over the scored items only, which are 48 easy, 60 original and 105 hard, because the 18 unscored items fall in the second and third files.
Figure 4: Left: every cyclic rotation scored on its own, against the full-rotation average, on four (model, task) cells. On all four the average falls at or below the best single rotation, by 0.3 to 2.3 points, and above the worst by 3.5 to 9.3 points. Right: where the budget stops on one 18-option cell, over 600 decisions, for a rule with no minimum-rotation floor; the shipped default reads two rotations before it may stop, which moves the leftmost bar.
Model
Task
K
Threshold
Verify rate
Verify bound
Clears 1%
Rotations
vs. K
Qwen2.5-7B
massive_route
18
10.25
0.0017
0.0079
yes
10.61
1.70 ×
Qwen2.5-7B
newsgroups
20
9.50
0.0017
0.0079
yes
9.01
2.22 ×
Qwen3-8B
massive_route
18
—
—
—
—
—
—
Qwen3-8B
newsgroups
20
9.00
0.0033
0.0105
no
5.31
3.77 ×
Appendix
Table 10: The certificate with the threshold selected on one third of an unlabelled batch and the bound computed on the other two thirds, for that one threshold. The verify bound carries no selection multiplicity. Two of the four cells clear a 1% target this way; on massive_route with Qwen3-8B no threshold on the grid cleared it on the selection split.
Stopping statistic
Cells certified
Mean rotations saved
log-odds margin alone
4 of 4
3.46 ×
log-odds margin + unanimity
4 of 4
2.91 ×
probability gap alone
2 of 4
3.19 ×
probability gap + unanimity
2 of 4
2.64 ×
Appendix
Table 11: Which stopping statistic satisfies Eq. ( 5 ) at ε=0.01 , over four (model, task) cells. A rule with fewer certified cells did not run worse: no threshold on the grid cleared the target on those cells at any budget.
Stopping statistic
Model
Task
K
Threshold
Rotations
vs. K
Disagreement
log-odds margin alone
Qwen2.5-7B
massive_route
18
4.5
6.12
2.94 ×
0.0083
Qwen2.5-7B
newsgroups
20
6.0
6.00
3.34 ×
0.0000
Qwen3-8B
massive_route
18
9.5
5.37
3.36 ×
0.0000
Qwen3-8B
newsgroups
20
8.25
4.75
4.21 ×
0.0033
log-odds margin + unanimity
Qwen2.5-7B
massive_route
18
4.25
7.96
2.26 ×
0.0083
Qwen2.5-7B
newsgroups
20
6.0
6.63
3.02 ×
0.0000
Appendix
Table 12: Every cell each stopping rule certifies at ε=0.01 , with its threshold and the rotations it then reads. Rules with two rows did not certify the other two cells at any threshold on the grid. These thresholds come from a different calibration sample than the serving run of Section 7 , which is why the same cell appears there at 6.0.
Model
Blocks
Depth at 0.98
Depth at 0.95
Before crossing
At crossing
Agreement at full depth
Mistral-7B
32
66%
66%
0.320 at 59%
0.985 at 66%
1.000
Qwen3-8B
36
89%
69%
0.087 at 64%
0.965 at 69%
0.993
Qwen3-1.7B
28
100%
79%
0.188 at 75%
0.958 at 79%
0.990
Qwen2.5-7B
28
100%
89%
0.420 at 79%
0.950 at 89%
0.983
Qwen3-32B
64
100%
91%
0.228 at 80%
0.973 at 91%
0.990
gpt-oss-20b
24
100%
92%
0.207 at 79%
0.950 at 92%
0.983
Appendix
Table 13: Agreement with the model’s own full-depth decision, for a readout taken at a fraction of the blocks and mapped into the final basis. Before crossing and at crossing are the last depth measured below 0.90 agreement and the first at or above it.
Figure 5: Agreement against depth, eight models. Markers are the depths we ran, and the curve is drawn through them rather than smoothed, because the intermediate depths were not measured. Hollow markers show the first depth reaching 0.95. Dashed lines mark 0.95 and 0.98.
System / readout
Overall ↑
95% CI
Choice
Yes/no
Ordinal
ECE ↓
n=18,808
n=17,572
n=9,940
(shard mean)
Reference rules (not learned systems)
Uniform random (expected)
34.29
—
25.77
50.00
22.63
—
Majority by question (test oracle)
54.28
—
41.42
73.57
44.51
—
Published baseline interfaces
Bespoke-Nimble-9B
70.57
[70.16, 70.99]
73.60
81.06
46.33
0.0889
Appendix
Table 14: All full-split BEV Decision Mix results. Accuracy and its intervals are in percent; ECE is on a 0–1 scale. Higher accuracy and lower ECE are better. Bold marks the best measured system in each metric, including ties.