cs.LGOct 1, 2026

Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning

Authors: Zhanming Zhang, Vinoth Selvendran

Organizations: Independent Researcher New York, NY · Independent Researcher Palo Alto, CA

Abstract

Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides noisier self-supervision and may spend much of a limited adaptation budget merely concentrating probability mass before reliably expressing capability it already possesses. We propose \textbf{entry-state sharpening}: use data-free training \emph{before} TTRL to prepare a general-purpose checkpoint in a state that subsequent label-free adaptation can exploit more efficiently. The idea is not tied to one training recipe; different data-free objectives can move the same base model to different entry states. Across five data-free checkpoints derived from Qwen3-4B and evaluated under an identical 15-step TTRL protocol, entry policy entropy strongly rank-orders endpoint conversion efficiency, a reliability-to-reachability measure (Spearman ρ=−0.90ρ=-0.90; ρ=−0.99ρ=-0.99 after controlling for entry reachability). The contrast across objectives is striking: R-Zero remains diffuse at 3.393.39 nats and finishes below the untuned base in 6/6 matched comparisons across MATH, GPQA, and AMC, whereas SPIRAL reaches 0.070.07 nats and achieves the highest post-TTRL accuracy on MATH and GPQA despite its self-play stage using no math training data. An in-domain label-free self-distillation intervention further shows that the entry state can be deliberately sharpened. These results motivate treating checkpoint preparation as a \emph{state-control problem}: use data-free training to improve TTRL readiness, with entry entropy as a label-free control signal and reachable capability as the constraint.

Figures & tables

Explore similar work

CardsList
  1. Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

    Sep 16, 2026Naveen Vakada, Mingyuan Li, Shaoxiong JiStable Test-Time AdaptationTest-Time Adaptation

  2. Label-Free Reinforcement Learning via Cross-Model Entropy

    May 27, 2026Matt Gorbett, Hossein ShiraziReinforcement Learning Post-TrainingCross-Entropy Losses

  3. TTSR: Test-Time Self-Evolving via Reflection

    Feb 6, 2026Haoyang He, Zihua Rong, Yunjia Zhao +3Test-Time TrainingTest Time