Prefill-only decision models inspired by the Jev model score every candidate in a menu during a single forward pass and never decode, which makes one call one to two orders of magnitude cheaper than a same-scale generative language model. We show that this read-out structure comes with a testable property. When an intervention changes only the candidate menu and leaves the input text fixed, the post-intervention accuracy is already determined by the cached first-pass distribution. The estimator restricts the pass-1 probabilities to the menu, renormalizes, and reads off the argmax; it uses no labels and no second forward pass. Across seven model families, ten datasets and two task types, menu-only interventions are predicted to within 4.2 points, and for one family the prediction is exact. A probability-level variant of the same estimator errs by 21.0 points, so the property lives in the ranking rather than in the probabilities and is not recovered by calibration. Same-scale generative language models do not share the property. On those models the same estimator errs by 1.6 to 15.8 points and degrades as the model grows. The property turns inference-time compute into a decision that can be made before deployment. Uniform extra passes buy calibration but almost no accuracy; at matched cost a confidence cascade outperforms every scheme that re-asks the same model, and curating the menu beats enlarging the model, with a 0.8B model on a curated 5-candidate menu reaching 95.4% on CLINC150 against 80.0% for a 4B model on the full 150-label menu.Code and data are available at https://github.com/rlisml/jev-cascade.
Figures & tables
Figure 1: Menu construction moves the accuracy across 26 points and the offline prediction matches it (left); escalation outperforms re-asking at matched cost (right).
Figure 2: The two read-out structures (top) and the three threads of the paper (bottom).
Family
rand 15
avoid 15
self clean
self verify
self 2
kev-0.8B
+0.5
+1.7
+6.8
−0.8
+9.8
OJ-2B
+0.0
+0.0
+0.0
+9.4
−0.4
dec-0.8B
−0.8
−0.6
−3.4
−3.0
−2.8
dec-2B
−0.4
−0.4
−2.0
−2.6
−3.6
this-that
+0.6
+3.2
−0.4
−1.8
−8.2
Mica-4B
+2.2
+0.2
+8.2
+8.8
+5.2
Table 1: Residual (pp) by model family and menu type. Bold marks an exact prediction; positive entries are overestimates.
Source
0.8B
4B
9B
match-9B
best
best cost
AG News
88.8
89.8
89.2
1.04
91.0
1.63
DBpedia
99.2
99.4
99.4
n/a
99.2
1.02
SciQ
95.9
99.0
99.2
1.80
99.4
2.07
ARC
61.6
92.2
94.4
5.88
94.8
6.82
CSQA
58.0
77.6
78.2
5.43
78.8
5.89
OBQA
50.6
84.0
85.8
6.04
87.2
6.83
Table 2: Three-tier kev cascade (0.8B → 4B → 9B). Costs are in 0.8B forward units; match-9B is the cheapest cascade reaching 9B accuracy. Bold marks the cascade optimum.
Figure 3: Readout stability. From left to right, the panels show prediction against measurement, residual by intervention family, label-free baselines, and the same estimator applied to generative models.
Figure 4: Where the property breaks. From left to right, the panels show residual against confusability, residual against menu size, position on a small menu, and a per-item postdiction of the hardest exception.
Figure 5: The economics of decision-time compute. From left to right, the panels show the accuracy gain from sixteen order-permuted passes, calibration error against K , the equal-budget anchor, and gate quality across tasks.
Figure 6: Candidate-set governance and presentation. From left to right, the panels show the TREC hierarchy collapse, avoidance against random menus, offline menu search, and nine phrasings of one menu.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Residual type
Size (pp)
Mechanism
pure restriction
∣Δ∣≤2
context-free menu changes; (P) holds
re-ask anchoring
+6 to +20
presentation artefact; removable
re-ask narrowing
−3 to −4
sign flips with label-space confusability
de-biasing excess
−1 to −4
permutation marginalisation averages out order noise
composed collapse
+11 to +53
two-stage decision multiplies the pass-1 bias
calibration drag
−9 to +3
cascade inherits the raw confidence bias
Appendix
Table 3: Residual typology.
Strategy
Cost (u)
Accuracy
flat, no re-ask
1.00
0.644
always re-ask (verify)
2.00
0.636
marginalisation K=2
2.00
0.652
SARP gate (verify)
1.44
0.668
gate, standard wording
1.29
0.664
cascade, escalate 5.6%
1.28
0.684
Appendix
Table 4: Same-budget comparison on 250 held-out CLINC150 items; costs are in 0.8B forward units.
Source
Model
Menu
hit rate
acc ∣ hit
r∗
verdict
GoEmotions
jev2b
avoid 15
0.322
0.426
0.90
deployable
GoEmotions
jev2b
random 15
0.553
0.453
0.85
deployable
MASSIVE
jev2b
avoid 15
0.091
0.713
0.79
deployable
MASSIVE
jev2b
random 15
0.269
0.775
0.73
deployable
GoEmotions
kev0.8b
avoid 15
0.280
0.368
1.01
never beats flat
GoEmotions
kev0.8b
random 15
0.553
0.441
0.85
deployable
Appendix
Table 5: Deployment with a menu that excludes the gold label; the system beats flat when recall is at least r∗ .
Label-table branch
re-decide
map (no re-run)
prediction
auto-merge
67.20
71.40
70.60
oracle merge
68.60
69.20
69.35
random control
61.98
68.05
67.53
flat 150 baseline
66.25
Appendix
Table 6: Label-table governance on 2,000 held-out CLINC150 items.
Figure 7: The deployment pipeline.
Figure 8: Additional results. From left to right, the panels show CLINC150 menu construction, the cross-family sign flip, the presentation ladder, and label-table repair.