cs.LGMay 2, 2025

Focus on Likely Classes for Test-Time Prediction

Authors: Johannes Schneider

Organizations: Department of Computer Science and Information Systems University of Liechtenstein Vaduz, Liechtenstein

Abstract

We ask: Can focusing on likely classes of a single, in-domain sample improve accuracy? Prior work argued "no", we answer "yes" on average. Standard entropy minimization yields largest gains, and we further dissect why by looking at its two core mechanisms: increasing likely and decreasing unlikely class predictions. Simply maximizing logits of the two most likely classes explains most of its benefits and, in some cases, even outperforms entropy minimization. Thus, decreasing classes is a secondary concern. Our controlled experiment and theory investigate the most puzzling case, why simply optimizing the logits of the two most likely classes leads to accuracy gains. Our small scale networks suggest that the degree of uncertainty of a prediction, followed by model choice, and to a lesser extent dataset choice moderate the success chances of the optimization. We show that the true class tends to have larger gradients independent of whether the prediction is correct or not. Our theory thoroughly investigates one of multiple possible explanations from the angle of shared features. Our evaluation focuses on 26 large language models and 12 text datasets, and 18 image recognition models trained on ImageNet. It demonstrates gains of 0.6 percent on average for entropy minimization on uncertain samples using no extra data aside from the given test sample using a fixed learning rate. We also suggest gradient-ray ensembling. When traversing input/embedding space by applying the same gradient with differing learning rates and aggregating probabilities for the same sample further improves accuracy significantly to 0.75 percent relying on a coarse range for learning rates, thus improving accuracy and eliminating the need to find a precise learning rate at the expense of computation.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Apr 28, 2026cs.LG

Entropy Centroids as Intrinsic Rewards for Test-Time Scaling

An effective way to scale up test-time compute of large language models is to sample multiple responses and then select the best one, as in Grok Heavy and Gemini Deep Think. Existing selection methods often rely on external reward models, which requires training a strong reward model and introduces additional computation overhead. As an alternative, previous approaches have explored intrinsic signals, such as confidence and entropy, but these signals are noisy with naive aggregation. In this work, we observe that high-entropy tokens tend to cluster into consecutive groups during inference, providing a more stable notion of model uncertainty than individual tokens. Together, these clusters reveal temporal patterns of model uncertainty throughout the inference process. Motivated by this observation, we propose to use the temporal structure of uncertainty as an intrinsic reward. To this end, we first formalize the basic unit of segment-level uncertainty as the High Entropy Phase (HEP), a variable-length segment that begins at a high-entropy token and ends when consecutive low-entropy tokens appear. We then define the Entropy Centroid, inspired by the concept of the center of mass in physics, as the weighted average position of all HEPs along the trajectory. Intuitively, a lower centroid indicates early exploration followed by confident generation, which we find often corresponds to higher response quality. Based on this insight, we propose the Lowest Centroid method, which selects the response with the lowest entropy centroid among multiple candidates. Experiments on mathematics, code generation, logical reasoning, and agentic tasks, across model scales ranging from 14B to 480B, show that Lowest Centroid consistently outperforms existing baselines and delivers stable gains as model size increases. Code is available at https://github.com/hkust-nlp/entropy-centroid.
Jun 19, 2026cs.LG

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning

Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expensive training, curated data, and reward signals. Recent work shows that sampling from sharpened base-model distributions at test time recovers much of the RL gain, yet existing methods rely solely on output-layer likelihoods and ignore the transformer's internal forward-pass dynamics. We introduce Depth-Entropy Guided Sampling (DEGS), a training-free, test-time method that exploits layer-wise entropy collapse as an intrinsic quality signal. We observe that stronger reasoners -- including RL-posttrained variants -- exhibit a distinctive "late collapse": logit-lens decoded entropy stays elevated until deeper layers before converging. We define a per-sequence collapse depth D(x)D(\mathbf{x}) and a joint objective π(x)∝p(x)αexp⁡(βD(x))π(\mathbf{x}) \propto p(\mathbf{x})^α\exp(βD(\mathbf{x})) that combines sequence likelihood with this depth-entropy structure, instantiated inside an MCMC power-sampling framework (DEGS-MCMC). Across three open-weight models and four reasoning benchmarks, this near-chance per-candidate signal compounds over the sampling trajectory into state-of-the-art training-free accuracy, with gains largest out of domain and on the harder splits -- exactly where likelihood alone falls short -- at single-digit-percent wall-clock overhead. DEGS narrowly trails an in-house GRPO reference on the math splits GRPO was trained for, yet surpasses it out of domain on GPQA for all three models, without any training, reward model, or labeled data.
Jun 1, 2026cs.LG

Entropy Minimization without Model Collapse: Mitigating Prediction Bias in Medical Imaging

Entropy minimization (EM) is the dominant objective for test-time adaptation, yet its failure mode, model collapse, remains poorly understood. In this work, we show that distribution shifts can cause feature clusters corresponding to distinct classes in the model's representation space to merge, while the decision boundary remains fixed. This induces a systematic skew in the predicted class distribution, referred to as prediction bias. Prediction bias refers to a shift in the predicted class distribution, with some classes overrepresented and others suppressed. We show that entropy minimization amplifies this prediction bias by tightening the existing clusters, reinforcing the incorrect groupings until all predictions collapse to a trivial solution. Next, to demonstrate the significance of prediction bias and mitigate it, we further propose Distribution Shift Bias Reduction (DSBR), a bias-correcting objective that specifically targets this failure mode by equalizing the contribution of each predicted class to the unsupervised entropy minimization loss. To study this failure mode, we design suitable adaptation settings using four medical-imaging datasets and additionally evaluate on ImageNet-C. We find that DSBR consistently stabilizes test-time adaptation, prevents model collapse, and matches or outperforms state-of-the-art methods. Moreover, DSBR operates solely at test-time.