Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries
Authors: Chonghe Jiang, Ao Qu, Siyuan Liu, Ruoyun Ma, Zijian Zhou, Dingyi Zhuang, Bo Liu, Han Zheng, +4 more
Organizations: Massachusetts Institute of Technology · Singapore-MIT Alliance for Research and Technology · Hong Kong Polytechnic University · ByteDance Inc. · National University of Singapore · Stanford University · University of Washington · University of California, Berkeley · Tsinghua University
Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Methods such as TTT-Discover use test-time training (TTT) to update the solution-generating LLM from verifier feedback, adapting its generation policy to improve subsequent proposals on the target problem. However, this becomes expensive when reliable execution requires a large model, since training must maintain gradients, optimizer states, and policy statistics while repeatedly generating long, structured outputs. It also complicates credit assignment: outcome-level verifier feedback must jointly evaluate the high-level strategy and its low-level implementation. In this work, we introduce Guidance-TTT, which separates these roles. A compact guidance model is trained at test time to propose high-level strategic changes, while a frozen execution model implements them as complete executable solutions. At each step, the system selects a promising previously discovered solution, proposes a change, executes and verifies it, and updates only the guidance model using an adaptive group-relative RL objective. This concentrates test-time learning on short strategic decisions while retaining the implementation capability of a substantially stronger model without adapting it. Without web access, Guidance-TTT produces strong solutions across four distinct domains: combinatorial optimization (Polyomino Packing), heuristic programming (AHC058), machine learning (Lasso), and GPU kernel optimization (TriMul). Across these tasks, it outperforms the best solutions reported in prior work while remaining competitive with state-of-the-art results on public online leaderboards. Code is available at https://github.com/Human-Agent-Society/reef/tree/guidance-ttt-support.
Figures & tables
Polyomino ( Qu et al., 2026 )
Lasso ( Ye et al., 2026 )
AHC058 ( Ye et al., 2026 )
TriMul 1 ( Cao et al., 2026 )
Method
Score ↑
Score ↑
AtCoder score ↑
Latency ( μ s) ↓
Previous SOTA 2
89.40
0.1243
849,325,750
1131
Ours
91.89
0.1739
850,082,731
1129
Table 1: Discovery results across four domains.
Figure 1: Test-time learning in plan space. TTT-Discover uses verifier reward to update both plan and code generation. Guidance-TTT applies the reward only to the small guidance model, while a frozen executor translates plan changes into code changes.
Figure 2: Discovery paradigms differ in where verifier feedback drives adaptation: search in self-evolution, trainable solution generation in TTT-Discover, and trainable strategic guidance in Guidance-TTT. Training only the compact guidance model reduces adaptation cost while retaining the implementation capabilities of a stronger, frozen executor.
Figure 3: Guidance-TTT workflow. A PUCT-based archive search selects a promising parent solution, the guidance model proposes a high-level modification, and the frozen executor realizes it as a new candidate. The resulting verifier reward updates only the guidance model.
Method / solution
Model(s)
Result
Polyomino Packing Score ↑ ; 70 cases
CORAL ( Qu et al., 2026 )
Claude Opus 4.6
84.20
CORAL
GLM-5.2
83.80
TTT-Discover ( Yuksekgonul et al., 2026 )
GPT-OSS-120B
83.72
OpenEvolve ( Sharma, 2025 )
GLM-5.2
78.36
Human best ( Mang et al., 2025 )
N/A
89.10
Table 2: Comparison with prior work and public references across four discovery tasks. Bold marks the best displayed value per task.
Figure 4: Comparison of solution structures. Blue denotes prior approaches; gold denotes ours.
Setting
Polyomino score ↑
TriMul latency ( μ s) ↓
Matched 8B configurations
Guidance-TTT
91.89
1225.58
Frozen guidance
84.85
1256.62
Direct solution training
41.63
9664.91
Extra parent context (G-Top2)
82.12
1166.82
Table 3: Ablations on Polyomino and TriMul. Bold marks the best result among the matched 8B configurations.
Figure 5: TriMul model-capability probe. Counts report explanations without incorrect extra conditions across three samples per model (higher is better).
Figure 6: Best-so-far search trajectories under GLM-5.2 execution. Search scores and Table 2 results use different evaluation environments.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Real pass
Real time ↓
Stress pass
Stress time ↓
All OOD pass
SimpleTES public best ( Ye et al., 2026 )
4/4
3.370
9/11
–
13/15
Guidance-TTT, Qwen3-8B + GLM-5.2
4/4
4.003
11/11
8.068
15/15
Guidance-TTT, Qwen3-8B + GPT-OSS-120B
2/4
–
11/11
11.828
13/15
Guidance-TTT, Qwen3-14B + GLM-5.2
4/4
3.479
11/11
8.535
15/15
Appendix
Table 4: Frozen Lasso OOD results. Time is the geometric mean in milliseconds; a dash indicates that at least one case failed correctness.
Latency (ms) ↓
Speedup ↑
Dataset
SimpleTES
Ours, 8B
Ours, 14B
8B
14B
Gisette
1,850.35 †
597.56
2,111.72
–
–
RCV1
10,225.08
10,156.79
63,674.64
1.01 ×
0.16 ×
DNA
10.16
7.50
13.68
1.35 ×
0.74 ×
Leukemia
10.32
9.34
8.05
1.10 ×
1.28 ×
Colon Cancer
7.45
4.76
4.45
1.57 ×
1.68 ×
Appendix
Table 5: Lasso transfer to 11 real OOD datasets. Latency is the mean of five runs in milliseconds ( ↓ ); speedup is SimpleTES latency divided by ours ( ↑ ). Both solvers were discovered using GLM-5.2 as the executor. Bold marks the fastest valid solver in each row.
Method / solution
Triton 3.3.1 development ( μ s)
Triton 3.4.0 OOD ( μ s)
Triton 3.6.0 OOD ( μ s)
Automated discovery systems
Guidance-TTT, Qwen3-14B guidance (ours)
1129
1124
1072
K-Search ( Cao et al., 2026 )
1131
1169
1154
TTT-Discover ( Yuksekgonul et al., 2026 )
1262
1229
1164
Aster ( Bicker, 2026 )
1280
1232
1212
Public GPUMode submissions
Appendix
Table 6: TriMul geometric-mean latency for fixed kernels across Triton versions. Our kernel is selected on Triton 3.3.1 and evaluated unchanged on the OOD versions. GPUMode ranks follow the public leaderboard snapshot cited in this paper ( GPU MODE, 2026 ) . Underlining marks the fastest automated discovery result in each column.
Executor
Updates
Polyomino ↑
GLM-5.2
30
91.89
GPT-5.4
30
85.32
GPT-OSS-120B
30
83.19
Appendix
Table 7: Polyomino executor variants with Qwen3-8B guidance.
Figure 7: Relative change from 8B to 14B guidance with GLM-5.2 execution. Positive values indicate improvement; TriMul shows relative latency reduction, and AHC058 compares best training-archive scores. TriMul’s 6.4% reduction uses the same-batch means of 1297.44 μ s (8B) and 1214.62 μ s (14B). Each size has one search trajectory. Raw validity includes service failures, which are audited separately below. Absolute statistics are in Appendix B.1.1 .
Task
8B guidance
14B guidance
Polyomino ↑
91.89
89.80
Lasso ↑
0.1739±0.0015
0.1730±0.0023
AHC058 ↑
849,063,233
850,082,731
TriMul ( μ s) ↓
1225.58
1128.91
Appendix
Table 8: Retained results by guidance size across four tasks.
Task
Guidance size
Validity (%)
Useful yield (%)
Polyomino
8B
89.48
14.09
14B
93.05
13.59
32B
83.15
9.30
TriMul
8B
63.80
8.70
14B
68.96
9.53
Lasso
8B
86.48
7.21
Appendix
Table 9: Absolute candidate statistics for GLM-5.2 execution. Validity and useful yield use all 3,840 candidate attempts as denominator, including service failures. Useful candidates must pass the verifier and strictly improve their exact selected parent.
Figure 8: Both guidance sizes on Lasso and AHC058, with the same GLM-5.2 seed within each task. These are search-time archive scores. Lasso’s 14B archive peak (0.23467) exceeds the 8B peak (0.21259), but its frozen-solver mean (0.1730) does not exceed the 8B mean (0.1739).
Task
Check
8B
14B
Polyomino
Score of an exact fit
1/3
2/3
TriMul
Condition for skipping a whole tile
2/3
3/3
AHC058
Inclusion of exact phase boundaries
2/3
3/3
AHC058
Reachability of a joint upgrade
1/3
3/3
Appendix
Table 10: Selected checks where 14B answers correctly more often than 8B. Counts are correct responses out of three to the same prompt; larger is better. For the AHC058 boundary check, both boundaries must be correct.
Figure 9: Two failure modes in Guidance-TTT. (a) Execution failures can assign zero reward to useful guidance. (b) The archive–policy loop can concentrate training and search within a few successful lineages.
Figure 10: Cases corresponding to the two failure modes in Figure 9 . (a) Two proposals from the same Polyomino parent receive different learning signals because one execution fails. (b) The step-best candidates at updates 28–30 add concavity-aware gap signatures, density-adaptive residual-flexibility weights, and complementarity-based decomposition feedback, respectively. They retain the update-20 solver’s Bottom-left-fill skyline packing and simulated-annealing search and do not exceed the update-27 frontier.
Figure 11: TriMul cases corresponding to the two failure modes in Figure 9 . (a) Two proposals from the same parent use the same large- N persistent-BMM schedule, but only one realization passes correctness tests. (b) After the global frontier stops improving at update 25, the step-best candidates at updates 28–30 continue to refine BMM tiling and dispatch within the same kernel family.
Figure 12: Lasso examples of the same two failure structures studied in Polyomino and TriMul. (a) Related alignment proposals from one parent have different outcomes because one implementation contains a pointer-conversion error. (b) During a seven-update plateau, all selected parents share an update-9 ancestor; the displayed step-best programs refine the retained LARS/CD solver without exceeding the update-19 frontier. Values are search-time scores, not fixed-program re-evaluation means.
Figure 13: AHC058 counterparts of the Polyomino/TriMul failure cases. (a) Related proposals request adaptive investment phases and lookahead; one execution fails on a declaration-order error while another passes the judge. (b) During updates 23–26, selected parents share an update-18 ancestor and useful local changes do not improve the frontier. Displayed scores are means over 150 public instances, in millions, rather than the totals used in the main results. Both panels use the 14B-guidance run.
Component
Polyomino
TriMul
Guidance model
Qwen/Qwen3-8B , LoRA rank 32
Qwen/Qwen3-14B , LoRA rank 32; 8B for ablations
Execution model
z-ai/glm-5.2 , OpenRouter, high reasoning
z-ai/glm-5.2 , OpenRouter, high reasoning
Hardware
1 B200 for guidance RL; hosted execution
1 B200 for guidance RL; isolated H100 verifiers
Rollout shape
8 parents/update × 16 rollouts/parent
Sampling
Guidance temperature 1.0, top- p=1.0 ; provider high-reasoning execution
Guidance temperature 1.0, top- p=1.0 ; high-reasoning execution
Actor limits
4,096 prompt; 8,192 response tokens
6,144 prompt; 10,240 response tokens
Appendix
Table 11: Primary configurations for Polyomino Packing and TriMul. The Polyomino execution API had no client-side completion cap. Every Guidance-TTT result on TriMul uses the 8×16 rollout shape.
Total raw score over 150 public instances; 2s per instance
Same three configurations
Executor-generated one-shot seed
Appendix
Table 12: Configurations for the two additional tasks. Runs use eight parents per update and 16 rollouts per parent, with a 30-update budget. The AHC058 120B-executor run stopped after 28 updates.
Test-time training (TTT) adapts large language models (LLMs) during inference using only unlabeled test inputs. Existing methods, however, face two major bottlenecks on hard reasoning tasks: (1) \emph{lack of learnable samples}, as self-generated pseudo-labels on difficult questions are often noisy and yield unstable rewards; and (2) \emph{inefficient exploration}, as performance gains depend on repeatedly sampling many rollouts without explicit diagnosis of why previous attempts fail. We propose \textbf{TTSR} (\textbf{T}est-\textbf{T}ime \textbf{S}elf-\textbf{R}eflection), a self-evolving framework based on a \emph{reflect-then-synthesize} paradigm. A single pretrained model alternates between a \textit{Student} role and a \textit{Teacher} role: the Student solves test questions and updates, while the Teacher analyzes failed trajectories and synthesizes targeted variant questions closer to the Student's capability frontier. TTSR further maintains a cross-iteration \textit{weakness memory} and compiles persistent weaknesses into a lightweight \textit{strategy note} prepended to subsequent Student inputs, so diagnostic knowledge can guide exploration and gradually fade as weaknesses are resolved. Experiments on challenging mathematical reasoning benchmarks show consistent test-time improvements, strong cross-backbone generalization, and transfer to general-domain reasoning tasks.
Haoyang He, Zihua Rong, Yunjia Zhao +3
Beijing University of Posts and Telecommunications · Southwestern University of Finance and Economics · China Unicom Online Information Technology Co., Ltd.
Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrated into evolutionary search have recently produced state-of-the-art solutions on optimization tasks, including open mathematical conjectures, GPU kernel design, scientific law discovery, and combinatorial puzzles. To achieve this, prior work applied search scaffolds to one target task at a time, so every new problem is approached from scratch and the experience accumulated during search is discarded once the model finishes its attempt. This leaves the capability of iteratively evolving a solution (e.g., knowing which part to mutate and how, deciding when to backtrack) entirely in the scaffold rather than in the model itself. Whether the model itself could acquire this capability and reuse it across different tasks has been largely unexamined. To address this, we introduce Evolution Fine-Tuning (EFT), a mid-training paradigm that teaches LLMs to evolve solutions across tasks by converting evolutionary search trajectories into supervision. We construct Finch Collection, a 156K-trajectory dataset spanning 10 domains and 371 optimization tasks, and fine-tune open-source LLMs from 2B to 9B parameters. Empirically, EFT confers cross-task generalization: across 22 held-out tasks, our models surpass their base counterparts by 10.22% on average. Furthermore, when paired with test-time RL, our model matches state-of-the-art performance on two circle-packing tasks and outperforms its base-model counterpart on the Erdős minimum-overlap problem. EFT thus serves as a "practice phase" for general-purpose discovery agents that do not solve new problems from scratch.
Young-Jun Lee, Seungone Kim, Minki Kang +5
University of Minnesota · Carnegie Mellon University · KAIST +3
Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during inference. However, existing TTS strategies are largely hand-crafted: researchers manually design reasoning patterns and tune heuristics by intuition, leaving much of the computation-allocation space unexplored. We propose an environment-driven framework, AutoTTS, that changes what researchers design: from individual TTS heuristics to environments where TTS strategies can be discovered automatically. The key to AutoTTS lies in environment construction: the discovery environment must make the control space tractable and provide cheap, frequent feedback for TTS search. As a concrete instantiation, we formulate width--depth TTS as controller synthesis over pre-collected reasoning trajectories and probe signals, where controllers decide when to branch, continue, probe, prune, or stop and can be evaluated cheaply without repeated LLM calls. We further introduce beta parameterization to make the search tractable and fine-grained execution trace feedback to improve discovery efficiency by helping the agent diagnose why a TTS program fails. Experiments on mathematical reasoning benchmarks show that the discovered strategies improve the overall accuracy--cost tradeoff over strong manually designed baselines. The discovered strategies generalize to held-out benchmarks and model scales, while the entire discovery costs only $39.9 and 160 minutes. Our data, and code will be open-source at https://github.com/zhengkid/AutoTTS.