While modern large language models (LLMs) have been trained to reason through verbalized chains-of-thought, the generation cost grows substantially due to suboptimal paths to reach the final answer. Furthermore, as new insights are discovered while observing various input queries (e.g. through self-reflection), limited mechanisms exist for carrying forward these findings to be applied to subsequent problems. One can view the list of such strategies or behaviors as a growing cheatsheet, with elements retrieved from this memory module at inference-time. In this work, we consider structured cheatsheets, with learned clusters of behaviors. We introduce a Hierarchical Dirichlet Process Gaussian Mixture Model (HDP-GMM) over behavior embeddings, which shares components across domains while allowing domain-specific mixing weights, and uses the posterior predictive to retrieve relevant behaviors for a query; we call this a Bayesian Cheatsheet. This mechanism allows for cheap adaptation in an online test-time training (TTT) setting, softly updating the mixture's sufficient statistics following each sample and enabling the creation of new components when the synthesized behaviors are sufficiently novel. We demonstrate that Bayesian Cheatsheet achieves clear performance gains relative to existing memory modules across reasoning benchmarks such as AIME'25, Omni-MATH, and PhysReason, even in the cold-start setting. We show that the Bayesian Cheatsheet is an adaptively reorganizing memory module, as behaviors can be re-assigned to components through a single step of collapsed Gibbs sampling. Our findings highlight the value of Bayesian-inspired memory modules for effective test-time adaptation and the role of structure in metacognitive reasoning.
Figures & tables
Figure 2: Static retrieval with Qwen3-32B using the mined behavior corpus. Bayesian Cheatsheet – specifically, our posterior predictive strategy (BC-PP) – achieves the strongest performance across the board, while our top- k′ nearest-centroid variant also performs well relative to the baselines.
Qwen3-32B
Granite-4.2-30B
Method
Condition
AIME’25
OM
PR
AIME’25
OM
PR
Hard k -means
Static
73.3
20.3
18.7
75.6
27.0
18.3
Online, no refit
75.6
24.5
23.5
77.8
28.6
20.4
Online, with refit
74.4
24.8
24.8
76.7
31.6
24.6
Soft k -means
Static
73.3
21.6
23.6
75.6
28.4
23.2
Online, no refit
75.6
23.4
28.5
78.9
26.9
22.0
Table 1: Static prompting vs. online adaptation, with and without periodic refitting ( g=50 for OM and PR, g=10 for AIME’25), for Qwen3-32B and Granite-4.2-30B. The static version of DC-RS is Metacognitive Reuse (in Table 4 ). DC-Cu and DC-RS involve "curation" where the cheatsheet can be refined (does not have a separate refit), while ACE similarly uses a curator for adaptation and does not have a static equivalent. We denote these entries with a "–". Bayesian Cheatsheet (BC-TTT) improves over existing methods across models and benchmarks, both with and without refit.
Figure 3: Assignment churn (Qwen3-32B) on PhysReason-mini; refits every 10 queries.
Qwen3-32B
Granite-4.2-30B
Method
OM
PR
OM
PR
Base model (no memory)
18.1
16.4
24.8
16.4
DC-Cu, cold-start
22.5
15.9
25.5
15.3
ACE, cold-start
21.0
17.7
27.9
16.4
BC-TTT, cold-start (ours)
24.1
20.5
31.9
18.1
Table 2: Accuracy on Omni-MATH ( n=916 ) and PhysReason ( n=226 ) with no pre-mined behavior corpus. Memory-based methods start from an empty memory and accumulate behaviors through self-reflection at test time. Bayesian Cheatsheets use a periodic refit interval of g=50 .
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Global concentration γ
5.0
Domain concentration αd
1.0
NIW scale Ψ0
1.2× empirical covariance in PCA space
NIW degrees of freedom ν0
D+2=130
NIW mean strength κ0
0.01
Projection dimension D
128
Appendix
Table 3: Hyperparameters for HDP-GMM fitting and retrieval, as well as the refit window. The refit settings only apply to the settings where periodic refitting is used.
Qwen3-32B
Granite-4.2-30B
Method
AIME’25
OM
PR
AIME’25
OM
PR
Base model
71.1
18.1
16.4
73.3
24.8
16.4
Random behavior selection
72.2
18.9
18.6
73.3
27.5
16.4
Direct behavior retrieval
72.2
18.6
23.5
74.4
28.2
19.5
Metacognitive Reuse, k′=5
75.6
17.8
21.7
77.8
28.4
22.6
BC-centroids (ours)
74.4
20.5
24.8
81.1
29.1
23.9
Appendix
Table 4: Static-retrieval results using the mined behavior corpus for Qwen3-32B and Granite-4.2-30B. OM and PR denote Omni-MATH and PhysReason, respectively. Bold and underlined values indicate the best and second-best results within each model and benchmark.
Table 5: Effect of the retrieval size k′ (number of components) on static BC-PP accuracy.
g
Accuracy
Kfinal
Net increase
Novelty-triggered
0
26.4
76
12
12
5
24.8
94
30
5
10
25.2
86
22
4
25
25.2
85
21
4
50
26.6
85
21
11
75
26.4
85
21
14
Appendix
Table 7: Effect of the refit interval g on online Qwen3-32B accuracy on Omni-MATH. K is the final component count (initial is 64 ), “net increase” is the net change in the number of components over the run, and “novelty” denotes the subset of those new components that were novelty-triggered.
Mixture structure
Condition
AIME’25
Omni-MATH
PhysReason
HDP-GMM (hierarchical)
Online, no refit
81.1
26.4
30.1
DP-GMM (pooled)
Online, no refit
76.7
24.0
27.3
Appendix
Table 8: Hierarchical (HDP-GMM) vs. flat, domain-pooled (DP-GMM) mixture structure for Qwen3-32B in the online setting without refitting (all other mechanisms and hyperparameters are held fixed). The hierarchical structure clearly improves performance on all three benchmarks.
Configuration
Omni-MATH
Threshold, KL-merge
26.6
Sampling, KL-merge
26.0
Threshold, no KL-merge
25.9
Sampling, no KL-merge
26.2
Appendix
Table 9: Performance on Omni-MATH under four combinations of new-component creation mechanism (threshold vs. sampling) and KL-based component consolidation (with vs. without). All experiments involve periodic refitting ( g=50 ).
Protocol
Kfinal
Net increase
Threshold
76
+12 (novelty-triggered)
Sampling
138
+74 (Bernoulli-triggered)
Appendix
Table 10: Component count and change on Omni-MATH without periodic refitting ( g=0 ), comparing threshold vs. sampling-based component creation mechanisms. All component births are dependent on the component creation mechanism.
Configuration
Kfinal
Net increase
Online-spawned
Refit-triggered (Net)
Threshold, KL-merge
85
+21
11
10
Sampling, KL-merge
129
+65
69
−4
Threshold, no KL-merge
82
+18
2
16
Sampling, no KL-merge
130
+66
68
−2
Appendix
Table 11: Component counts and changes on Omni-MATH with periodic refitting ( g=50 ), comparing threshold vs. sampling-based component creation mechanisms.
τnew
Accuracy
Kfinal
Net increase
Novelty-triggered
Refit-triggered (Net)
0.1
24.3
206
+142
212
−70
0.3
26.0
127
+63
65
−2
0.5
26.6
85
+21
11
+10
0.7
26.1
84
+20
0
+20
0.9
26.0
83
+19
0
+19
Appendix
Table 12: Comparing the impact of the novelty threshold on Omni-MATH performance and component growth; all experiments use the main configuration (threshold-based component creation, refitting with g=50 , and KL-based component merging during refitting).
#
Domain
Label
0
Math
Unit consistency & dimensional-analysis checks
1
Physics
Multi-method cross-checks & literature verification
2
Math
Solution verification & domain-restriction checks
3
Physics
Cosmology & dark-matter physics
4
Math
Symmetry-based algebraic simplification
5
Physics
Applied materials & optics engineering formulas
Appendix
Table 13: Short thematic labels for each of the K=64 HDP-GMM components fit on the Qwen3-32B mined behavior corpus.
#
Domain
Label
0
Physics
Chaos theory & dynamical-systems criteria
1
Math
Trigonometric identities in optics & triangle geometry
Table 14: Short thematic labels for each of the K=62 components in the HDP-GMM memory module fit from the Granite-4.2-30B-mined behavior corpus.
#
Domain
Representative behavior
0
Math
Always check and convert units to ensure consistency before performing calculations, using known conversion factors if necessary.
4
Math
When solving equations, check for symmetry in variables (e.g., if swapping x and y yields the same equation) to simplify or identify solutions.
8
Physics
Use symmetry to simplify field contributions: opposite sides of a loop may cancel if their contributions are equal and opposite, but verify directions to avoid assuming cancellation.
14
Math
When solving for area, cross-check the result with perimeter calculations (if relevant) to ensure logical consistency.
18
Physics
Use multiple independent sources or methods to confirm the value of critical parameters (e.g., particle masses) before finalizing a solution.
30
Physics
When analyzing perturbations, always check if they break a symmetry that protects the system’s properties.
Appendix
Table 15: Selected representative behaviors from a subset of the K=64 HDP-GMM components fit from the Qwen3-32B-mined behavior corpus (component numbers match those in Table 13 )
Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience across them can further improve performance. Existing memory-based methods either store whole-solution templates that generalize poorly to novel problems or use heuristic step-level selection that is not optimized for final-answer correctness. Learning selection policies requires large-scale training data and fixed action spaces, making such approaches unsuitable for test-time settings where memory expands incrementally and only limited supervision is available. We propose MILES (Modular Instruction Memory with LEarnable Selection for self-improving LLM reasoning), a framework that dynamically expands step-wise memory and applies correctness-optimized memory composition under realistic test-time constraints. MILES maintains modular memory units consisting of asymmetric pairs of sub-goal embeddings and sub-instructions, each associated with a learnable selection head. This memory structure enables a coarse-to-fine retrieval mechanism: The coarse level enables memory expansion and collects supervision for training selection heads from confident samples, while the fine stage applies learned selection heads to rerank coarse-level candidates and guide reasoning for uncertain samples. MILES consistently matches or outperforms prior methods while achieving superior accuracy-efficiency tradeoffs. Extensive experiments demonstrate its effectiveness, robustness, and transferability.
Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it. We formalize test-time reasoning as a constrained compute allocation problem over partial trajectories. Under a fixed hardware budget, existing paradigms fail to actively allocate the compute to the most promising partial progress: traditional parallel sampling treats traces independently and induces severe memory bottlenecks, while subtractive pruning starves hardware and fails to actively and sufficiently shift the output distribution. To overcome this dichotomy, we introduce Gambit, an inference algorithm that executes \emph{thought-level beam search}. By periodically pruning unpromising trajectories and immediately branching from high-quality prefixes, Gambit dynamically concentrates compute onto the most promising reasoning traces via a light-weight scorer probing hidden states while maintaining continuous high hardware utilization. Extensive evaluations across multiple models and benchmarks demonstrate that Gambit strictly dominates existing baselines. Under identical hardware constraints, our method yields up to a +6.7% absolute accuracy gain on HMMT-24 and +3.3% on AIME-25 over pruning baselines, delivers >2× higher throughput on trace completion, and reduces total token consumption by up to 68.5% relative to standard parallel sampling.
Large language models (LLMs) are typically deployed with fixed parameters, and their performance is often improved by allocating more computation at inference time. While such test-time scaling can be effective, it cannot correct model misconceptions or adapt the model to the specific structure of an individual query. Test-time optimization addresses this limitation by enabling parameter updates during inference, but existing approaches either rely on external data or optimize generic self-supervised objectives that lack query-specific alignment. In this work, we propose Query-Conditioned Test-Time Self-Training (QueST), a framework that adapts model parameters during inference using supervision derived directly from the input query. Our key insight is that the input query itself encodes latent signals sufficient for constructing structurally related problem--solution pairs. Based on this, QueST generates such query-conditioned pairs and uses them as supervision for parameter-efficient fine-tuning at test time. The adapted model is then used to produce the final answer, enabling query-specific adaptation without any external data. Across seven mathematical reasoning benchmarks and the GPQA-Diamond scientific reasoning benchmark, QueST consistently outperforms strong test-time optimization baselines. These results demonstrate that query-conditioned self-training is an effective and practical paradigm for test-time adaptation in LLMs. Code is available at https://chssong.github.io/Query-Conditioned-TTST/.
Chaehee Song, Minseok Seo, Yeeun Seong +2
School of Electrical Engineering, KAIST · 2Graduate School of Green Growth and Sustainability, KAIST