Probability-only models, which TypeSafe calls System One models, return calibrated probabilities for fixed choices in milliseconds and generate no text. We study one such model, Jev, through two tasks that require decisions under tight constraints. In bullet chess, a bot that places Jev's judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots. Live model calls are too slow for search, so we distill pairwise judgments into a compact evaluator that runs at every position. We then ask how best to spend a fixed labeling budget when an LLM, Qwen3-32B, is available as a second teacher. In chess, averaging both judges' labels beats spending the whole budget on Qwen alone by 9.6 Elo (95% interval 4.3 to 14.9), and the gain replicates on fresh openings; a second answer from the same judge is no substitute, and Jev is the strongest partner for Qwen among the models tested. In passage reranking, Jev's labels alone train a reranker that scores as high as Qwen's, from 21 minutes of API calls instead of 5.1 GPU-hours, and adding Qwen gains at most a few thousandths in ranking quality. Search supplies the lookahead, distillation makes the judgment cheap enough to use at every position, and an LLM partner pays off in chess.
Figures & tables
Figure 1: The complete bot’s Lichess bullet trajectory at one minute per side. The strip groups wins (green), draws (grey), and losses (orange) by deployment round.
Figure 2: (a) Distilling judge probabilities into search weights. (b) Combined evaluator vs. Qwen on twice the pairs at equal budget; 2,000, 2,000, 4,000 (pooled), and 4,000 games. (c) Substitute controls, 4,000 games each. Dots: estimates; lines: 95% intervals.
Rival (252,540 responses each)
Elo
95% interval
Jev, twice the pairs
+27.0
[ 19.2 , 34.8 ]
Qwen, twice the pairs
+8.0
[ 0.9 , 15.3 ]
Qwen, two answers averaged
+14.3
[ 5.9 , 22.5 ]
Jev, two answers averaged
+28.2
[ 20.2 , 36.3 ]
Qwen, adjusted toward 0.5 ( τ=0.4465 )
+8.0
[ 0.5 , 15.4 ]
Table 1: Jev+Qwen vs. each rival at 252,540 responses, 2,000 games. Row 2 clusters by single opening; rows 1 and 4 by 20-ply prefix; rows 3 and 5 by 8-ply.
Substitute for Jev
Elo
95% interval
Jev replaced by 0.5
+9.6
[ 3.7 , 15.6 ]
Jev probabilities shuffled freely
+10.9
[ 5.3 , 16.8 ]
Qwen alone, recalibrated, twice the pairs
+14.3
[ 8.8 , 20.2 ]
Jev shuffled within Qwen groups
+5.3
[ −0.5 , 11.0 ]
Table 2: Jev+Qwen vs. four substitutes, 4,000 games on fresh openings at the same response budget.
Jev+Qwen’s lead over
Elo
95% interval
Full-data settings
OLMo+Qwen
+16.6
[ 10.3 , 22.8 ]
DiffusionGemma+Qwen
+16.9
[ 10.2 , 23.4 ]
Settings selected per subset
OLMo+Qwen
+7.7
[ 1.4 , 14.0 ]
DiffusionGemma+Qwen
+7.9
[ 1.3 , 14.6 ]
Table 3: Jev+Qwen’s lead over alternative Qwen partnerships, 6,000 games per evaluator across three subsets. Intervals resample openings.
Figure 3: Jev+Qwen’s lead over alternative partnerships, 6,000 games per evaluator. Filled: full-data settings; open: per-subset. Lines: 95% intervals.
Figure 4: (a) Passage-reranking NDCG@10 on TREC DL 2019+2020. Open dots: individual training runs; tick: mean over five seeds, with both order assignments pooled for the split row. (b) Agreement with human graders on 500 pairs.
Judge
ms/resp
Timing
Hardware
Cost
Jev
7.6
20 concurrent
API; no local GPU
$1.97
Qwen3-32B
54.6
sequential
1 × H200
1.91 GPU-h
OLMo-2-32B
35.6
sequential
PRO 6000 + H200
1.25 GPU-h
DiffusionGemma
7.9
32 concurrent
1 × RTX 5090
0.28 GPU-h
Table 4: Cost of obtaining 126,270 judge responses (63,135 pairs, both orders). Effective rates reflect concurrent requests.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluator and scale c
Elo
95% interval
Jev, c=0.5
+32.2
[ 13.9 , 50.7 ]
Jev, c=0.125
+27.9
[ 10.4 , 46.3 ]
Qwen, c=0.5
+19.1
[ 4.3 , 34.0 ]
Qwen, c=0.125
+34.9
[ 16.5 , 53.4 ]
Appendix
Table 5: Jev+Qwen against single-judge evaluators trained on twice the pairs at the listed scales, 400 games on 200 unused openings.
Same pairs
Twice the pairs
c
Jev+Qwen
Jev
Qwen
Jev
Qwen
0.125
+33.1
+10.4
+24.4
+10.4
+29.6
0.25
+31.4
+10.4
0.0
+24.4
+34.9
0.5
+57.9
+12.2
+17.4
−3.5
+17.4
1.0
−5.2
−24.4
−125.0
−24.4
−68.6
Appendix
Table 6: Validation Elo against published tables. Each coefficient is the ratio of the learned evaluator’s standard deviation to the baseline’s. Bold marks the selected scale.
Judges
n
Elo
95% interval
Jev+Qwen
2
+25.8
[ 12.9 , 38.4 ]
Qwen
1
+24.0
[ 13.2 , 34.9 ]
Jev+Qwen+DiffusionGemma
3
+19.1
[ 8.0 , 30.3 ]
Qwen+OLMo+DiffusionGemma
3
+13.6
[ 2.1 , 25.1 ]
Jev
1
+13.2
[ 2.1 , 24.0 ]
OLMo+Qwen
2
+13.2
[ 2.4 , 24.7 ]
Appendix
Table 7: All 15 judge combinations, 1,000 games each against the published-table evaluator on shared openings. n : number of judges.
Weight group
Δ Elo
95% interval
Jev+Qwen, replacing from DiffusionGemma+Qwen
Piece-sq. ( 0.33× scale)
−7.7
[ −23.8 , 8.4 ]
Piece-sq. (matched scale)
−8.7
[ −25.8 , 8.4 ]
mobility
+3.1
[ −13.6 , 19.9 ]
king safety
−10.1
[ −27.2 , 6.6 ]
pawn structure
−7.3
[ −24.1 , 9.7 ]
Appendix
Table 8: Elo change from replacing one weight group, 1,000 games per row. Piece-square swaps change total score variation; rescaled swaps hold it fixed.
Figure 5: (a) Jev+Qwen vs. single-judge evaluators on twice the pairs at three time controls. (b) Changing one judge at equal budget, 2,000 games each. Dots: estimates; lines: 95% intervals.
Table 9: Move accuracy (%) and playing strength. The complete evaluator includes material, mobility and the integer calculations used during play.
Evaluator
Subset 1
Subset 2
Subset 3
Jev+Qwen
+14.8
+7.8
+9.4
OLMo+Qwen
−10.3
−4.7
−2.8
DiffusionGemma+Qwen
−12.7
−2.3
−3.8
Qwen + Jev shuffled within Qwen groups
−1.6
+3.8
+4.2
Appendix
Table 10: Elo for each evaluator against the corresponding Qwen evaluator trained on 2 × the pairs, with 2,000 games per comparison. Each evaluator retains its full-data weight-size setting.