cs.AIOct 5, 2026

Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries

Authors: Chonghe Jiang, Ao Qu, Siyuan Liu, Ruoyun Ma, Zijian Zhou, Dingyi Zhuang, Bo Liu, Han Zheng, +4 more

Organizations: Massachusetts Institute of Technology · Singapore-MIT Alliance for Research and Technology · Hong Kong Polytechnic University · ByteDance Inc. · National University of Singapore · Stanford University · University of Washington · University of California, Berkeley · Tsinghua University

Abstract

Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Methods such as TTT-Discover use test-time training (TTT) to update the solution-generating LLM from verifier feedback, adapting its generation policy to improve subsequent proposals on the target problem. However, this becomes expensive when reliable execution requires a large model, since training must maintain gradients, optimizer states, and policy statistics while repeatedly generating long, structured outputs. It also complicates credit assignment: outcome-level verifier feedback must jointly evaluate the high-level strategy and its low-level implementation. In this work, we introduce Guidance-TTT, which separates these roles. A compact guidance model is trained at test time to propose high-level strategic changes, while a frozen execution model implements them as complete executable solutions. At each step, the system selects a promising previously discovered solution, proposes a change, executes and verifies it, and updates only the guidance model using an adaptive group-relative RL objective. This concentrates test-time learning on short strategic decisions while retaining the implementation capability of a substantially stronger model without adapting it. Without web access, Guidance-TTT produces strong solutions across four distinct domains: combinatorial optimization (Polyomino Packing), heuristic programming (AHC058), machine learning (Lasso), and GPU kernel optimization (TriMul). Across these tasks, it outperforms the best solutions reported in prior work while remaining competitive with state-of-the-art results on public online leaderboards. Code is available at https://github.com/Human-Agent-Society/reef/tree/guidance-ttt-support.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TTSR: Test-Time Self-Evolving via Reflection

    Feb 6, 2026Haoyang He, Zihua Rong, Yunjia Zhao +3LLM Self-RefinementTest-Time Optimization

  2. Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks

    Jun 27, 2026Young-Jun Lee, Seungone Kim, Minki Kang +5Supervised Fine-TuningEvolutionary Optimization

  3. LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

    May 8, 2026Tong Zheng, Haolin Liu, Chengsong Huang +10LLM Inference EfficiencyAutomated Algorithm Discovery