Prefill-only decision models inspired by the Jev model score every candidate in a menu during a single forward pass and never decode, which makes one call one to two orders of magnitude cheaper than a same-scale generative language model. We show that this read-out structure comes with a testable property. When an intervention changes only the candidate menu and leaves the input text fixed, the post-intervention accuracy is already determined by the cached first-pass distribution. The estimator restricts the pass-1 probabilities to the menu, renormalizes, and reads off the argmax; it uses no labels and no second forward pass. Across seven model families, ten datasets and two task types, menu-only interventions are predicted to within 4.2 points, and for one family the prediction is exact. A probability-level variant of the same estimator errs by 21.0 points, so the property lives in the ranking rather than in the probabilities and is not recovered by calibration. Same-scale generative language models do not share the property. On those models the same estimator errs by 1.6 to 15.8 points and degrades as the model grows. The property turns inference-time compute into a decision that can be made before deployment. Uniform extra passes buy calibration but almost no accuracy; at matched cost a confidence cascade outperforms every scheme that re-asks the same model, and curating the menu beats enlarging the model, with a 0.8B model on a curated 5-candidate menu reaching 95.4% on CLINC150 against 80.0% for a 4B model on the full 150-label menu.Code and data are available at https://github.com/rlisml/jev-cascade.
Figures & tables
Figure 1: Menu construction moves the accuracy across 26 points and the offline prediction matches it (left); escalation outperforms re-asking at matched cost (right).
Figure 2: The two read-out structures (top) and the three threads of the paper (bottom).
Family
rand 15
avoid 15
self clean
self verify
self 2
kev-0.8B
+0.5
+1.7
+6.8
−0.8
+9.8
OJ-2B
+0.0
+0.0
+0.0
+9.4
−0.4
dec-0.8B
−0.8
−0.6
−3.4
−3.0
−2.8
dec-2B
−0.4
−0.4
−2.0
−2.6
−3.6
this-that
+0.6
+3.2
−0.4
−1.8
−8.2
Mica-4B
+2.2
+0.2
+8.2
+8.8
+5.2
Table 1: Residual (pp) by model family and menu type. Bold marks an exact prediction; positive entries are overestimates.
Source
0.8B
4B
9B
match-9B
best
best cost
AG News
88.8
89.8
89.2
1.04
91.0
1.63
DBpedia
99.2
99.4
99.4
n/a
99.2
1.02
SciQ
95.9
99.0
99.2
1.80
99.4
2.07
ARC
61.6
92.2
94.4
5.88
94.8
6.82
CSQA
58.0
77.6
78.2
5.43
78.8
5.89
OBQA
50.6
84.0
85.8
6.04
87.2
6.83
Table 2: Three-tier kev cascade (0.8B → 4B → 9B). Costs are in 0.8B forward units; match-9B is the cheapest cascade reaching 9B accuracy. Bold marks the cascade optimum.
Figure 3: Readout stability. From left to right, the panels show prediction against measurement, residual by intervention family, label-free baselines, and the same estimator applied to generative models.
Figure 4: Where the property breaks. From left to right, the panels show residual against confusability, residual against menu size, position on a small menu, and a per-item postdiction of the hardest exception.
Figure 5: The economics of decision-time compute. From left to right, the panels show the accuracy gain from sixteen order-permuted passes, calibration error against K , the equal-budget anchor, and gate quality across tasks.
Figure 6: Candidate-set governance and presentation. From left to right, the panels show the TREC hierarchy collapse, avoidance against random menus, offline menu search, and nine phrasings of one menu.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Residual type
Size (pp)
Mechanism
pure restriction
∣Δ∣≤2
context-free menu changes; (P) holds
re-ask anchoring
+6 to +20
presentation artefact; removable
re-ask narrowing
−3 to −4
sign flips with label-space confusability
de-biasing excess
−1 to −4
permutation marginalisation averages out order noise
composed collapse
+11 to +53
two-stage decision multiplies the pass-1 bias
calibration drag
−9 to +3
cascade inherits the raw confidence bias
Appendix
Table 3: Residual typology.
Strategy
Cost (u)
Accuracy
flat, no re-ask
1.00
0.644
always re-ask (verify)
2.00
0.636
marginalisation K=2
2.00
0.652
SARP gate (verify)
1.44
0.668
gate, standard wording
1.29
0.664
cascade, escalate 5.6%
1.28
0.684
Appendix
Table 4: Same-budget comparison on 250 held-out CLINC150 items; costs are in 0.8B forward units.
Source
Model
Menu
hit rate
acc ∣ hit
r∗
verdict
GoEmotions
jev2b
avoid 15
0.322
0.426
0.90
deployable
GoEmotions
jev2b
random 15
0.553
0.453
0.85
deployable
MASSIVE
jev2b
avoid 15
0.091
0.713
0.79
deployable
MASSIVE
jev2b
random 15
0.269
0.775
0.73
deployable
GoEmotions
kev0.8b
avoid 15
0.280
0.368
1.01
never beats flat
GoEmotions
kev0.8b
random 15
0.553
0.441
0.85
deployable
Appendix
Table 5: Deployment with a menu that excludes the gold label; the system beats flat when recall is at least r∗ .
Label-table branch
re-decide
map (no re-run)
prediction
auto-merge
67.20
71.40
70.60
oracle merge
68.60
69.20
69.35
random control
61.98
68.05
67.53
flat 150 baseline
66.25
Appendix
Table 6: Label-table governance on 2,000 held-out CLINC150 items.
Figure 7: The deployment pipeline.
Figure 8: Additional results. From left to right, the panels show CLINC150 menu construction, the cross-family sign flip, the presentation ladder, and label-table repair.
A typed decision is a choice among a fixed set of options, returned as a probability rather than as text. Systems that need typed decisions today use models trained for that purpose. This report describes AnyJev, which reads a typed decision from one prefill of a pretrained instruction-tuned language model. The readout restricts the next-token distribution at the answer position to the option tokens. It has two defects: the model assigns higher probability to some labels whatever the input, and to some positions in the option list. AnyJev corrects both with no gradient steps and no parameter changes: it divides out a label prior estimated from unlabelled inputs, and it averages log-probabilities over the K cyclic rotations of the option list. On two 20-option tasks the rotations lower the order-flip rate from 0.33 to 0.14 and from 0.33 to 0.18, and raise accuracy on 11 of 11 models on both. Reading every rotation requires K prefills. A stopping rule selected against the full-rotation decision on unlabelled states cuts that. Selecting the threshold on one unlabelled split and bounding its disagreement on a second, it reads 10.6 rotations of 18 at a verified 0.008 bound on two of four cells; selected and bounded on one split, as our serving run did, it reads 7.3 and serves 2.2 times as many decisions per second on vLLM. The code is open source.
Jiamu Zhang, Tianze Yang, Yucheng Shi +6
Nokia, Sunnyvale, CA, USA · Tencent Hunyuan · The Hong Kong Polytechnic University
Dedicated decision models such as Jev map unstructured language to probability distributions over finite choices, allowing their outputs to directly route requests, select tools, and trigger actions. Yet real-world inputs rarely arrive in isolation: they come with background details and surrounding context. We find that short additions that fit naturally into this context can nevertheless redirect an otherwise correct decision, even when the correct answer remains unchanged. To study this behavior, we fix a wrong target option for each initially correct item and use the model's option probabilities to refine fluent context additions while preserving the source, question, choices, and gold answer. Within 64 accepted target evaluations, the optimizer identifies contexts that redirect Jev on 312 of 508 initially correct decisions (61.4%); in 229 cases, Jev assigns at least 0.7 probability to the fixed wrong option. Across seven datasets, three additional decision systems show targeted flip rates of 64.9%-73.2% on decisions they initially answer correctly. Taken together, these results expose a pronounced fragility in current decision models: short, ordinary-looking context can shift a correct choice to a high-confidence wrong one. Because these models turn language directly into downstream choices, this sensitivity raises concerns about treating their probability outputs as reliable decision interfaces.
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM-as-Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.