Probability-only models, which TypeSafe calls System One models, return calibrated probabilities for fixed choices in milliseconds and generate no text. We study one such model, Jev, through two tasks that require decisions under tight constraints. In bullet chess, a bot that places Jev's judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots. Live model calls are too slow for search, so we distill pairwise judgments into a compact evaluator that runs at every position. We then ask how best to spend a fixed labeling budget when an LLM, Qwen3-32B, is available as a second teacher. In chess, averaging both judges' labels beats spending the whole budget on Qwen alone by 9.6 Elo (95% interval 4.3 to 14.9), and the gain replicates on fresh openings; a second answer from the same judge is no substitute, and Jev is the strongest partner for Qwen among the models tested. In passage reranking, Jev's labels alone train a reranker that scores as high as Qwen's, from 21 minutes of API calls instead of 5.1 GPU-hours, and adding Qwen gains at most a few thousandths in ranking quality. Search supplies the lookahead, distillation makes the judgment cheap enough to use at every position, and an LLM partner pays off in chess.
Figures & tables
Figure 1: The complete bot’s Lichess bullet trajectory at one minute per side. The strip groups wins (green), draws (grey), and losses (orange) by deployment round.
Figure 2: (a) Distilling judge probabilities into search weights. (b) Combined evaluator vs. Qwen on twice the pairs at equal budget; 2,000, 2,000, 4,000 (pooled), and 4,000 games. (c) Substitute controls, 4,000 games each. Dots: estimates; lines: 95% intervals.
Rival (252,540 responses each)
Elo
95% interval
Jev, twice the pairs
+27.0
[ 19.2 , 34.8 ]
Qwen, twice the pairs
+8.0
[ 0.9 , 15.3 ]
Qwen, two answers averaged
+14.3
[ 5.9 , 22.5 ]
Jev, two answers averaged
+28.2
[ 20.2 , 36.3 ]
Qwen, adjusted toward 0.5 ( τ=0.4465 )
+8.0
[ 0.5 , 15.4 ]
Table 1: Jev+Qwen vs. each rival at 252,540 responses, 2,000 games. Row 2 clusters by single opening; rows 1 and 4 by 20-ply prefix; rows 3 and 5 by 8-ply.
Substitute for Jev
Elo
95% interval
Jev replaced by 0.5
+9.6
[ 3.7 , 15.6 ]
Jev probabilities shuffled freely
+10.9
[ 5.3 , 16.8 ]
Qwen alone, recalibrated, twice the pairs
+14.3
[ 8.8 , 20.2 ]
Jev shuffled within Qwen groups
+5.3
[ −0.5 , 11.0 ]
Table 2: Jev+Qwen vs. four substitutes, 4,000 games on fresh openings at the same response budget.
Jev+Qwen’s lead over
Elo
95% interval
Full-data settings
OLMo+Qwen
+16.6
[ 10.3 , 22.8 ]
DiffusionGemma+Qwen
+16.9
[ 10.2 , 23.4 ]
Settings selected per subset
OLMo+Qwen
+7.7
[ 1.4 , 14.0 ]
DiffusionGemma+Qwen
+7.9
[ 1.3 , 14.6 ]
Table 3: Jev+Qwen’s lead over alternative Qwen partnerships, 6,000 games per evaluator across three subsets. Intervals resample openings.
Figure 3: Jev+Qwen’s lead over alternative partnerships, 6,000 games per evaluator. Filled: full-data settings; open: per-subset. Lines: 95% intervals.
Figure 4: (a) Passage-reranking NDCG@10 on TREC DL 2019+2020. Open dots: individual training runs; tick: mean over five seeds, with both order assignments pooled for the split row. (b) Agreement with human graders on 500 pairs.
Judge
ms/resp
Timing
Hardware
Cost
Jev
7.6
20 concurrent
API; no local GPU
$1.97
Qwen3-32B
54.6
sequential
1 × H200
1.91 GPU-h
OLMo-2-32B
35.6
sequential
PRO 6000 + H200
1.25 GPU-h
DiffusionGemma
7.9
32 concurrent
1 × RTX 5090
0.28 GPU-h
Table 4: Cost of obtaining 126,270 judge responses (63,135 pairs, both orders). Effective rates reflect concurrent requests.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluator and scale c
Elo
95% interval
Jev, c=0.5
+32.2
[ 13.9 , 50.7 ]
Jev, c=0.125
+27.9
[ 10.4 , 46.3 ]
Qwen, c=0.5
+19.1
[ 4.3 , 34.0 ]
Qwen, c=0.125
+34.9
[ 16.5 , 53.4 ]
Appendix
Table 5: Jev+Qwen against single-judge evaluators trained on twice the pairs at the listed scales, 400 games on 200 unused openings.
Same pairs
Twice the pairs
c
Jev+Qwen
Jev
Qwen
Jev
Qwen
0.125
+33.1
+10.4
+24.4
+10.4
+29.6
0.25
+31.4
+10.4
0.0
+24.4
+34.9
0.5
+57.9
+12.2
+17.4
−3.5
+17.4
1.0
−5.2
−24.4
−125.0
−24.4
−68.6
Appendix
Table 6: Validation Elo against published tables. Each coefficient is the ratio of the learned evaluator’s standard deviation to the baseline’s. Bold marks the selected scale.
Judges
n
Elo
95% interval
Jev+Qwen
2
+25.8
[ 12.9 , 38.4 ]
Qwen
1
+24.0
[ 13.2 , 34.9 ]
Jev+Qwen+DiffusionGemma
3
+19.1
[ 8.0 , 30.3 ]
Qwen+OLMo+DiffusionGemma
3
+13.6
[ 2.1 , 25.1 ]
Jev
1
+13.2
[ 2.1 , 24.0 ]
OLMo+Qwen
2
+13.2
[ 2.4 , 24.7 ]
Appendix
Table 7: All 15 judge combinations, 1,000 games each against the published-table evaluator on shared openings. n : number of judges.
Weight group
Δ Elo
95% interval
Jev+Qwen, replacing from DiffusionGemma+Qwen
Piece-sq. ( 0.33× scale)
−7.7
[ −23.8 , 8.4 ]
Piece-sq. (matched scale)
−8.7
[ −25.8 , 8.4 ]
mobility
+3.1
[ −13.6 , 19.9 ]
king safety
−10.1
[ −27.2 , 6.6 ]
pawn structure
−7.3
[ −24.1 , 9.7 ]
Appendix
Table 8: Elo change from replacing one weight group, 1,000 games per row. Piece-square swaps change total score variation; rescaled swaps hold it fixed.
Figure 5: (a) Jev+Qwen vs. single-judge evaluators on twice the pairs at three time controls. (b) Changing one judge at equal budget, 2,000 games each. Dots: estimates; lines: 95% intervals.
Table 9: Move accuracy (%) and playing strength. The complete evaluator includes material, mobility and the integer calculations used during play.
Evaluator
Subset 1
Subset 2
Subset 3
Jev+Qwen
+14.8
+7.8
+9.4
OLMo+Qwen
−10.3
−4.7
−2.8
DiffusionGemma+Qwen
−12.7
−2.3
−3.8
Qwen + Jev shuffled within Qwen groups
−1.6
+3.8
+4.2
Appendix
Table 10: Elo for each evaluator against the corresponding Qwen evaluator trained on 2 × the pairs, with 2,000 games per comparison. Each evaluator retains its full-data weight-size setting.
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM-as-Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.