Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides noisier self-supervision and may spend much of a limited adaptation budget merely concentrating probability mass before reliably expressing capability it already possesses. We propose \textbf{entry-state sharpening}: use data-free training \emph{before} TTRL to prepare a general-purpose checkpoint in a state that subsequent label-free adaptation can exploit more efficiently. The idea is not tied to one training recipe; different data-free objectives can move the same base model to different entry states. Across five data-free checkpoints derived from Qwen3-4B and evaluated under an identical 15-step TTRL protocol, entry policy entropy strongly rank-orders endpoint conversion efficiency, a reliability-to-reachability measure (Spearman ρ=−0.90; ρ=−0.99 after controlling for entry reachability). The contrast across objectives is striking: R-Zero remains diffuse at 3.39 nats and finishes below the untuned base in 6/6 matched comparisons across MATH, GPQA, and AMC, whereas SPIRAL reaches 0.07 nats and achieves the highest post-TTRL accuracy on MATH and GPQA despite its self-play stage using no math training data. An in-domain label-free self-distillation intervention further shows that the entry state can be deliberately sharpened. These results motivate treating checkpoint preparation as a \emph{state-control problem}: use data-free training to improve TTRL readiness, with entry entropy as a label-free control signal and reachable capability as the constraint.
Figures & tables
Figure 1: Identical 15-step TTRL on unlabeled MATH problems for Base, R-Zero, and SPIRAL@400; means over seeds {2,3} . R-Zero starts 4.5 accuracy points ahead of Base and finishes 5.8 points behind it, while SPIRAL@400 finishes highest despite never training on math.
Checkpoint
Objective
Entry
Entry
Entry
Final
Final
η
entropy
mean@16
best@16
mean@16
best@16
R-Zero
uncert. curriculum
3.391
0.717
0.799
0.713±0.004
0.783±0.007
0.911±0.013
Base
— (pretrained)
1.186
0.672
0.749
0.771±0.005
0.820±0.006
0.941±0.001
SPIRAL@32
game play, Kuhn @32
0.344
0.747
0.806
0.771±0.015
0.823±0.013
0.937±0.003
Self-distilled
data-free, high-dose
0.183
0.710
0.771
0.767±0.016
0.811±0.018
0.945±0.001
SPIRAL@400
game play, multi-game
0.073
0.755
0.796
0.787±0.005
0.822±0.001
0.957±0.006
Table 1: The five data-free checkpoints under the frozen 15-step TTRL protocol on MATH, ordered by entry entropy. Entry columns are seed means; final columns and η are means ± sample s.d. over seeds {2,3} . Endpoint conversion efficiency η is strongly ordered by entry entropy.
Figure 2: Entry policy entropy versus endpoint conversion efficiency for the five data-free checkpoints. Filled markers denote seed means and hollow markers individual seeds {2,3} .
Checkpoint
Objective
Entry
Entry
Entry
Final
Final
η
entropy
mean@16
best@16
mean@16
best@16
R-Zero
uncert. curriculum
4.644
0.305
0.731
0.317±0.006
0.735±0.024
0.432±0.021
Base
— (pretrained)
1.791
0.239
0.736
0.357±0.013
0.766±0.017
0.466±0.007
SPIRAL@32
game play, Kuhn @32
0.681
0.349
0.790
0.383±0.004
0.783±0.011
0.489±0.012
Self-distilled
data-free, high-dose
0.454
0.272
0.758
0.369±0.001
0.727±0.005
0.508±0.002
SPIRAL@400
game play, multi-game
0.130
0.383
0.780
0.406±0.002
0.755±0.010
0.538±0.005
Table 2: GPQA under the frozen 15-step TTRL protocol, matched seeds {2,3} , ordered by entry entropy. Entry columns are seed means; final columns and η are means ± sample s.d.
Checkpoint
Objective
Entry
Entry
Entry
Final
Final
η
entropy
mean@16
best@16
mean@16
best@16
R-Zero
uncert. curriculum
3.452
0.417
0.710
0.414±0.003
0.665±0.024
0.623±0.018
Base
— (pretrained)
1.041
0.393
0.739
0.460±0.006
0.747±0.001
0.615±0.007
SPIRAL@32
game play, Kuhn @32
0.278
0.426
0.684
0.460±0.013
0.725±0.004
0.635±0.021
Self-distilled
data-free, high-dose
0.221
0.403
0.741
0.479±0.006
0.700±0.002
0.684±0.011
SPIRAL@400
game play, multi-game
0.076
0.436
0.744
0.476±0.011
0.704±0.009
0.676±0.007
Table 3: AMC under the frozen 15-step TTRL protocol, matched seeds {2,3} , ordered by entry entropy. Entry columns are seed means; final columns and η are means ± sample s.d.
Beijing University of Posts and Telecommunications · Southwestern University of Finance and Economics · China Unicom Online Information Technology Co., Ltd.