Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoning. Existing approaches largely prescribe these decisions through their harness; we instead shift them to the model. We introduce 1) Hermes, a family of simple, configurable harnesses that progressively varies model control over context allocation and reuse, and 2) Hermes-Learn, a two-stage framework for learning these capabilities. We find that capable models can exploit this flexibility to scale with additional inference-time compute, while smaller open-source models initially struggle to do so. Training with Hermes-Learn closes this gap, inducing adaptive contextual reasoning strategies that vary with both the problem and the progress of reasoning. These gains generalize across benchmarks and models, extrapolate beyond the inference-time compute seen during training, and transfer to complementary test-time scaling methods beyond Hermes.
Figures & tables
Method
Model-driven strategy selection
Context compaction
Recursive delegation
SPIRAL
✗
✗
✗
PaCoRe
✗
✓
✗
RAO
✗
✗
✓
Context-Folding
✗
✓ †
✗
Reasoning Cache
✗
✓
✗
Hermes
✓
✓
✓
Table 1: Comparison of Hermes with existing methods. Prior methods mostly prescribe contextual reasoning strategies for allocating and reusing contexts, then train models for that particular harness. Hermes-Learn instead trains models to select among different contextual reasoning strategies through the learning curriculum induced by the varying steering of the Hermes hierarchy. Model-driven strategy selection captures whether contextual reasoning strategies are decided by the model rather than by the harness; Context compaction , whether prior reasoning can be compressed and reused; Recursive delegation , whether further contexts can be allocated recursively through the same delegation interface.
Figure 1 : Overview of Hermes . Hermes enables contextual reasoning across multiple context windows through two mechanisms: delegation , which allocates fresh contexts to search or verification strategies, and digestion , which compacts accumulated progress for reuse in later contexts. The Hermes hierarchy ( L3 – L1 ) progressively shifts control over delegation decisions from the harness to the model. L3 prescribes the strategy; L2 prescribes a strategy category (search or verification) while allowing the model to select strategies within it; and L1 lets the model choose both as reasoning unfolds.
Pooled ( n=238 )
Synthetic ( n=88 )
Model
Inline
L1
L2
Inline
L1
L2
Qwen3-4B
37.1
16.2
15.1
50.7
26.8
26.0
Sonnet 5
55.3
76.7
76.5
87.0
94.7
96.2
DeepSeek-V3.2
56.2
58.4
57.4
78.1
86.2
83.7
Table 2 : Hermes enables test-time scaling for capable models. Avg@16 (%, ↑ ) under inline reasoning and Hermes L1 and L2 on the pooled and synthetic benchmarks. Inline reasoning measures mathematical reasoning within a single 8K-token context, while Hermes provides up to 56K tokens across contexts. Claude Sonnet 5 and DeepSeek-V3.2 match or improve over inline reasoning with Hermes ; Qwen3-4B degrades, exposing a contextual reasoning gap.
Figure 2 : Synthetic splitting example. Constructed by combining two independent AceReason questions into a single problem whose final answer depends on both, encouraging parallel delegation.
Figure 3 : Hermes-Learn makes strategy selection problem- and progress-dependent. Split is used more often on synthetic problems designed to benefit from decomposition, while the second delegation step increasingly favors verification or continues the first-step strategy.
Scaling Knob
Value
SFT
SFT + RL
Default
–
39.9
47.3
W
2
44.3
52.2
3
45.2
53.0
B
2
39.7
46.4
4
39.7
49.7
D
2
40.8
55.6
Table 5: Hermes-Learn generalizes across compute budgets, test-time scaling methods, and model families. Avg@16 (%, ↑ ) on the pooled benchmark. (a) We independently vary the number of context windows ( W ), subagents ( B ), and delegation depth ( D ) beyond the default Hermes configuration used for training and evaluation: T=3 , B=3 , D=1 , and W=1 . (b) We evaluate Recursive Self-Aggregation (RSA) and the DeepSeekMath (DSM) Agent, both alone and combined with L1 . (c) We apply the same training pipeline to Olmo-3-7B-Instruct . All experiments use the best performing Hermes-Learn model. Bold and underlined mark best and second-best results where applicable.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Harness
Strategy menu offered at a delegating step
Step 1
Step 2
L3
One named strategy
1
1
L2-SS
The three search strategies
3
3
L2-SV
The three search strategies at step 1, and the three verification strategies at step 2
3
3
L1
All six strategies, at every delegating step
6
6
Appendix
Table 6: The harnesses in the Hermes hierarchy. Each cell gives the number of strategies offered at that step; the instruction text is otherwise identical. Delegation is disabled at the final step.
Figure 4 : Synthetic problem examples. Colored letters identify intermediate values and subanswers that determine the final answer. [ … ] denotes problem text omitted for brevity.
Qwen3-4B-Instruct-2507
Olmo-3-7B-Instruct
Temperature
0.7
0.6
Top- p
0.8
0.95
Top- k
20
–
Min- p
0.0
–
Appendix
Table 7: Sampling parameters.
Teacher
Split
Same
Different
Verification
Multiple Search Strategies
Sonnet 5
21.9
43.9
63.5
99.3
27.6
DeepSeek-V3.2
26.0
70.4
26.6
95.1
22.6
Appendix
Table 8: Percentage (%) of teacher trajectories containing different strategies. Percentages may sum to more than 100% since a single trajectory can contain multiple strategies.
Teacher
Hermes Harness
Step 1 Search
Step 2 Search
Sonnet 5
L1
99.9
19.0
Sonnet 5
L2-SV
100.0
15.5
Sonnet 5
L2-SS
99.9
29.2
DeepSeek-V3.2
L1
99.5
35.4
DeepSeek-V3.2
L2-SV
99.6
28.9
DeepSeek-V3.2
L2-SS
99.9
53.2
Appendix
Table 9: Percentage (%) of teacher trajectories containing a search strategy in each step.
Supervised Fine-tuning
Learning rate
10−5
Learning schedule
Cosine with warmup ratio 0.1
Training batch size
32
Epochs
4
Reinforcement Learning
Algorithm
Dr. GRPO [ 31 ]
Appendix
Table 10: Training hyperparameters.
Figure 5 : (a) Hermes-Learn makes strategy selection increasingly problem-dependent. (b) Correct trajectories perform verification more frequently. (c) The second step primarily performs verification and otherwise tends to continue the first-step strategy. (d) Verification behavior depends on the first-step search strategy.
May 29, 2026·Kaiyu Huang, Xingyu Wang, Mingze Kong +6Test-Time Scaling
School of Computer Science and Technology, Tongji University · Shenzhen Research Institute of Big Data, The Chinese University of Hong Kong, Shenzhen · School of Data Science, The Chinese University of Hong Kong, Shenzhen +2