Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoning. Existing approaches largely prescribe these decisions through their harness; we instead shift them to the model. We introduce 1) Hermes, a family of simple, configurable harnesses that progressively varies model control over context allocation and reuse, and 2) Hermes-Learn, a two-stage framework for learning these capabilities. We find that capable models can exploit this flexibility to scale with additional inference-time compute, while smaller open-source models initially struggle to do so. Training with Hermes-Learn closes this gap, inducing adaptive contextual reasoning strategies that vary with both the problem and the progress of reasoning. These gains generalize across benchmarks and models, extrapolate beyond the inference-time compute seen during training, and transfer to complementary test-time scaling methods beyond Hermes.
Figures & tables
Method
Model-driven strategy selection
Context compaction
Recursive delegation
SPIRAL
✗
✗
✗
PaCoRe
✗
✓
✗
RAO
✗
✗
✓
Context-Folding
✗
✓ †
✗
Reasoning Cache
✗
✓
✗
Hermes
✓
✓
✓
Table 1: Comparison of Hermes with existing methods. Prior methods mostly prescribe contextual reasoning strategies for allocating and reusing contexts, then train models for that particular harness. Hermes-Learn instead trains models to select among different contextual reasoning strategies through the learning curriculum induced by the varying steering of the Hermes hierarchy. Model-driven strategy selection captures whether contextual reasoning strategies are decided by the model rather than by the harness; Context compaction , whether prior reasoning can be compressed and reused; Recursive delegation , whether further contexts can be allocated recursively through the same delegation interface.
Figure 1 : Overview of Hermes . Hermes enables contextual reasoning across multiple context windows through two mechanisms: delegation , which allocates fresh contexts to search or verification strategies, and digestion , which compacts accumulated progress for reuse in later contexts. The Hermes hierarchy ( L3 – L1 ) progressively shifts control over delegation decisions from the harness to the model. L3 prescribes the strategy; L2 prescribes a strategy category (search or verification) while allowing the model to select strategies within it; and L1 lets the model choose both as reasoning unfolds.
Pooled ( n=238 )
Synthetic ( n=88 )
Model
Inline
L1
L2
Inline
L1
L2
Qwen3-4B
37.1
16.2
15.1
50.7
26.8
26.0
Sonnet 5
55.3
76.7
76.5
87.0
94.7
96.2
DeepSeek-V3.2
56.2
58.4
57.4
78.1
86.2
83.7
Table 2 : Hermes enables test-time scaling for capable models. Avg@16 (%, ↑ ) under inline reasoning and Hermes L1 and L2 on the pooled and synthetic benchmarks. Inline reasoning measures mathematical reasoning within a single 8K-token context, while Hermes provides up to 56K tokens across contexts. Claude Sonnet 5 and DeepSeek-V3.2 match or improve over inline reasoning with Hermes ; Qwen3-4B degrades, exposing a contextual reasoning gap.
Figure 2 : Synthetic splitting example. Constructed by combining two independent AceReason questions into a single problem whose final answer depends on both, encouraging parallel delegation.
Figure 3 : Hermes-Learn makes strategy selection problem- and progress-dependent. Split is used more often on synthetic problems designed to benefit from decomposition, while the second delegation step increasingly favors verification or continues the first-step strategy.
Scaling Knob
Value
SFT
SFT + RL
Default
–
39.9
47.3
W
2
44.3
52.2
3
45.2
53.0
B
2
39.7
46.4
4
39.7
49.7
D
2
40.8
55.6
Table 5: Hermes-Learn generalizes across compute budgets, test-time scaling methods, and model families. Avg@16 (%, ↑ ) on the pooled benchmark. (a) We independently vary the number of context windows ( W ), subagents ( B ), and delegation depth ( D ) beyond the default Hermes configuration used for training and evaluation: T=3 , B=3 , D=1 , and W=1 . (b) We evaluate Recursive Self-Aggregation (RSA) and the DeepSeekMath (DSM) Agent, both alone and combined with L1 . (c) We apply the same training pipeline to Olmo-3-7B-Instruct . All experiments use the best performing Hermes-Learn model. Bold and underlined mark best and second-best results where applicable.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Harness
Strategy menu offered at a delegating step
Step 1
Step 2
L3
One named strategy
1
1
L2-SS
The three search strategies
3
3
L2-SV
The three search strategies at step 1, and the three verification strategies at step 2
3
3
L1
All six strategies, at every delegating step
6
6
Appendix
Table 6: The harnesses in the Hermes hierarchy. Each cell gives the number of strategies offered at that step; the instruction text is otherwise identical. Delegation is disabled at the final step.
Figure 4 : Synthetic problem examples. Colored letters identify intermediate values and subanswers that determine the final answer. [ … ] denotes problem text omitted for brevity.
Qwen3-4B-Instruct-2507
Olmo-3-7B-Instruct
Temperature
0.7
0.6
Top- p
0.8
0.95
Top- k
20
–
Min- p
0.0
–
Appendix
Table 7: Sampling parameters.
Teacher
Split
Same
Different
Verification
Multiple Search Strategies
Sonnet 5
21.9
43.9
63.5
99.3
27.6
DeepSeek-V3.2
26.0
70.4
26.6
95.1
22.6
Appendix
Table 8: Percentage (%) of teacher trajectories containing different strategies. Percentages may sum to more than 100% since a single trajectory can contain multiple strategies.
Teacher
Hermes Harness
Step 1 Search
Step 2 Search
Sonnet 5
L1
99.9
19.0
Sonnet 5
L2-SV
100.0
15.5
Sonnet 5
L2-SS
99.9
29.2
DeepSeek-V3.2
L1
99.5
35.4
DeepSeek-V3.2
L2-SV
99.6
28.9
DeepSeek-V3.2
L2-SS
99.9
53.2
Appendix
Table 9: Percentage (%) of teacher trajectories containing a search strategy in each step.
Supervised Fine-tuning
Learning rate
10−5
Learning schedule
Cosine with warmup ratio 0.1
Training batch size
32
Epochs
4
Reinforcement Learning
Algorithm
Dr. GRPO [ 31 ]
Appendix
Table 10: Training hyperparameters.
Figure 5 : (a) Hermes-Learn makes strategy selection increasingly problem-dependent. (b) Correct trajectories perform verification more frequently. (c) The second step primarily performs verification and otherwise tends to continue the first-step strategy. (d) Verification behavior depends on the first-step search strategy.
Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning. However, recent studies reveal that sequential test-time scaling often yields diminishing or even negative returns, as longer traces exhibit increased uncertainty, error compounding, and drift from the original problem. We propose ThinkRetrieve, a test-time scaling framework that augments the reasoning traces of LRMs with dynamically retrieved solved examples at each reasoning step. Given an external corpus of problems paired with step-by-step solutions, ThinkRetrieve retrieves relevant exemplars at each intermediate step and injects them directly into the thinking trace, providing the model with guidance on how to reason rather than merely what facts are relevant. Experiments across five reasoning models (1.5B--8B parameters) on GSM-8K, MATH-500, AIME 2025, and SciQ demonstrate that ThinkRetrieve consistently improves accuracy over standard test-time scaling, with relative gains of up to 60% on AIME 2025.
While scaling test-time compute can substantially improve model performance, existing approaches either rely on static compute allocation or sample from fixed generation distributions. In this work, we introduce a test-time compute allocation framework that jointly adapts where computation is spent and how generation is performed. Our method begins with a warm-up phase that identifies easy queries and assembles an initial pool of question-response pairs from the test set itself. An adaptive phase then concentrates further computation on unresolved queries while reshaping their generation distributions through evolving in-context demonstrations -- conditioning each generation on successful responses from semantically related queries rather than resampling from a fixed distribution. Experiments across math, coding, and reasoning benchmarks demonstrate that our approach consistently outperforms existing baselines while consuming substantially less inference-time compute.
Bowen Zuo, Dongruo Zhou, Yinglun Zhu
University of California, Riverside · 2Indiana University Bloomington
In real-world deployments of large language models (LLMs), balancing inference quality and computational cost has become a central challenge. Existing approaches tackle this trade-off along two largely independent dimensions: model routing, which switches among models of different scales to match request complexity, and test-time scaling (TTS), which adjusts inference-time compute within a fixed model for fine-grained control. However, this decoupled design introduces inherent limitations. Model routing yields coarse-grained, discrete performance changes due to the sparse set of model scales, while single-model TTS often encounters capacity ceilings and exhibits diminishing returns as compute increases. Moreover, treating the two mechanisms separately restricts adaptability in dynamic inference environments. To overcome these limitations, we introduce Unified Inference Scaling (UIS), which unifies model routing and TTS in a single optimization space. Building on this formulation, we propose UniScale, an online framework that models adaptive UIS as a contextual multi-armed bandit problem and learns inference policies via LinUCB. The framework incorporates efficiency-aware learning and cost modeling to ensure stable and scalable optimization over high-dimensional action spaces. Evaluation shows that UniScale effectively exploits the synergy in the UIS space to deliver a fine-grained and consistently better quality-cost trade-off across diverse, dynamic inference scenarios.
Kaiyu Huang, Xingyu Wang, Mingze Kong +6
School of Computer Science and Technology, Tongji University · Shenzhen Research Institute of Big Data, The Chinese University of Hong Kong, Shenzhen · School of Data Science, The Chinese University of Hong Kong, Shenzhen +2