Decision models such as Jev answer questions with probabilities, which are only useful if they are calibrated. Open-source reproductions rely on supervised fine-tuning plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models: the model samples a rationale, and we score the answer distribution it commits to afterwards with a strictly proper scoring rule. A variance identity shows that scoring the mixture of several samples rewards disagreeing rationales, and that RLVR is exactly this mixture objective without its diversity term. Optimized naively, the per-rationale objective either switches reasoning off or is drowned out by policy-gradient noise, which leads to a two-stage recipe: calibrate, then reinforce. With Qwen3-1.7B on two reasoning tasks (3 seeds, paired tests), RLCD matches or beats SFT, RFT/STaR and GRPO (each temperature-scaled) in accuracy and beats all of them in selective prediction; on GSM8K answer verification a single query decides \gvTwoCovFive% of the items at ≤5% error, versus \gvGrpoCovFive% for GRPO. When uncertainty comes from annotator disagreement, RLCD provably cannot beat cross-entropy. Code and results: https://github.com/ZimmyGao/openjev-rlcd.
Figures & tables
Figure 1: The System-One collapse (MMLU-Pro, per-rationale objective, trained end to end). Without a KL anchor the four rationales per question become identical and empty after ∼ 100 steps (left: variance of u across rationales) while the readout learns to hedge (right). With a KL anchor the model keeps reasoning and learns to hedge after it. The no-KL run was stopped at step 150, after the collapse.
GSM8K-Verify (reasoning-essential)
MMLU-Pro (knowledge-heavy)
Method
Acc ↑
Brier ↓
AURC ↓
Cov@5% ↑
T
Acc ↑
Brier ↓
AURC ↓
Cov@20% ↑
T
base (zero-shot reasoning)
84.1
0.211
0.073
14.9
9.1
38.9
0.786
0.516
0.0
10.9
SFT + TS (direct answer)
70.4 ± 0.2
0.382 ± 0.009
0.174 ± 0.010
6.2 ± 2.0
2.1
42.3 ± 0.4
0.711 ± 0.002
0.386 ± 0.005
15.6 ± 0.7
1.6
RFT/STaR + TS
89.0 ± 1.6
0.207 ± 0.020
0.103 ± 0.013
0.9 ± 0.7
7.3
32.7 ± 0.9
0.830 ± 0.004
0.619 ± 0.011
0.0 ± 0.0
17.6
RFT/STaR + KL + TS
91.4 ± 0.5
0.160 ± 0.006
0.075 ± 0.010
11.6 ± 8.1
6.6
40.9 ± 0.6
0.761 ± 0.006
0.493 ± 0.008
0.0 ± 0.0
11.0
GRPO + KL + TS
91.8 ± 0.6
0.150 ± 0.009
0.061 ± 0.006
19.4 ± 7.3
7.2
42.4 ± 0.5
0.753 ± 0.006
0.481 ± 0.015
0.0 ± 0.0
11.1
Table 1: Single-query decisions after temperature scaling (mean ± s.d. over 3 seeds). Cov@ ϵ : fraction of test items that can be decided automatically at error rate ≤ϵ (5% for GSM8K-Verify, 20% for MMLU-Pro, whose error rates are higher). T : temperature fitted on dev (1 = calibrated as trained). Best in bold.
Figure 2: Risk–coverage of single-query decisions (after temperature scaling, averaged over seeds). Dotted: the error budget used for Cov@ ϵ in Table 1 . Temperature scaling cannot repair the ranking of overconfident models (GRPO, RFT), whose error rate is flat in coverage.
GSM8K-Verify
MMLU-Pro
Stage-2 continuation (400 steps, same checkpoint)
Δ Acc
Δ Brier
Δ AURC
Δ Acc
Δ Brier
Δ AURC
readout-only continuation
−1.7 ∗∗
+0.028 ∗∗
+0.011 ∗
+0.5
−0.010
−0.011
anchor only (REINFORCE noise, no reward)
−2.6 ∗∗
+0.031 ∗∗
+0.016 ∗∗
−0.1
+0.002
+0.007
GRPO (correctness reward)
+1.8 ∗∗
−0.028 ∗∗
−0.012 ∗∗
−1.1
+0.026 ∗∗
+0.033 ∗∗
RLCD-RL (proper-score reward)
+1.8 ∗∗
−0.029 ∗∗
−0.013 ∗∗
+0.6
−0.009
−0.009
Table 2: Stage-2 continuations forked from the same stage-1 checkpoint (3 seeds), change relative to the checkpoint (single query, after TS). Anchor only: the same score-function noise and KL anchor as RLCD-RL but no outcome reward.
Figure 3: Stage-2 reward matters. Brier improvement over continuing to train the readout only, for the same 400-step budget from the same checkpoint (95% bootstrap intervals over test items, 3 seeds pooled).
ChaosNLI (100 annotators / item), Qwen3-1.7B
JSD to humans ↓
Majority acc ↑
seeds
rationales
Direct answer, CE on the label stream (SFT)
0.074
66.2
3
—
Mixture objective λ=1 , 16-rationale mixture
0.121
63.1
3
degenerate (hits length limit)
Mixture objective λ=1 , single rationale
0.342
55.4
3
degenerate
Mixture objective λ=1 + KL, mixture
0.149
67.8
1
healthy (48 tokens)
Per-trace objective λ=0 , no KL
0.075
63.0
1
empty (System-One)
Table 3: ChaosNLI (400 test items). The target is the distribution of 100 human labels; every method sees one fresh annotator label per visit.
REINFORCE weight c (from scratch)
Acc ↑
Brier ↓
AURC ↓
rationale tokens
0 (readout only)
46.3
0.676
0.323
182
0.1
45.0
0.677
0.327
155
0.3
44.6
0.694
0.352
148
1 (unbiased)
43.7
0.703
0.354
178
Table 4: Weight of the score-function term when training from scratch (MMLU-Pro, seed 17, per-rationale objective with KL anchor, single query after TS). More policy gradient is monotonically worse; this motivates separating the stages.