Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.
Figures & tables
Figure 1: Comparison of Existing Test-Time Scaling Paradigms and Personalized Test-Time Scaling. Existing methods primarily optimize accuracy against latency or inference cost separately, whereas Personalized TTS targets their joint satisfaction under user-specific requirements. PersonTTS amortizes user-conditioned controller discovery by reusing requirement-matched controllers and procedural guidance from prior searches.
Method
AIME24-25
AIME26
HMMT24
HMMT25
(Discovery)
(Held-out)
(Discovery)
(Held-out)
Source
Target
Source
Target
Source
Target
Source
Target
ASC
13.53
0.14
7.03
0.00
11.91
4.89
12.03
4.98
ESC
13.28
0.14
6.22
0.00
11.22
4.89
11.83
4.98
ParallelProbeSR
23.77
10.19
23.89
9.35
23.81
19.30
18.07
11.51
AutoTTS ( β=0.5 )
12.11
16.41
10.52
16.03
7.04
2.96
7.25
3.23
Table 1: JSR (%; higher is better) across problem and profile splits. Source and target profiles follow the Setup, and problem splits are labeled separately. PersonTTS variants use the policy with the highest JSR on discovery problems across initialization and all five rounds. PersonTTS uses both reuse mechanisms. “w/o reuse” removes both the Guide and warm-start, while the other “w/o” variants remove the named mechanism. Dashes denote unreported results. Bold marks the highest reported value in each column.
Figure 2: Best-observed JSR on discovery sets across rounds. Curves show the cumulative maximum JSR on the discovery sets after initialization and each of the five candidate rounds (R0–R4), averaged over target profiles on AIME24–25 and HMMT24. The four curves compare PersonTTS with the variants without warm-start, without the Guide, and without both mechanisms.
Method
Round
Total
R0
R1
R2
R3
R4
AIME
PersonTTS (w/o reuse)
18.67 ∣ 1.58
17.37 ∣ 2.29
14.56 ∣ 2.20
12.33 ∣ 1.94
11.70 ∣ 2.29
74.63 ∣ 10.30
PersonTTS (w/o Guide)
12.74 ∣ 1.61
13.72 ∣ 2.03
14.70 ∣ 2.36
12.85 ∣ 2.27
12.31 ∣ 2.37
66.33 ∣ 10.64
PersonTTS (w/o warm-start)
13.60 ∣ 1.24
7.05 ∣ 1.29
7.56 ∣ 1.46
6.13 ∣ 1.28
5.35 ∣ 1.37
39.69 ∣ 6.64
PersonTTS
8.49 ∣ 1.15
8.90 ∣ 1.51
8.49 ∣ 1.67
7.83 ∣ 1.57
6.70 ∣ 1.50
40.41 ∣ 7.40
Table 2: Agent-call elapsed time and cost during target-profile discovery on discovery problems. Each cell reports time (minutes) ∣ cost (USD), with lower values better for both metrics. Values are averaged over 20 target profiles per benchmark and round. Total gives the per-profile sum across the five rounds. Variant names follow Table 1 .
Figure 3: Discovery-to-held-out generalization across rounds. Solid and dashed lines report the JSR of the policy produced at each discovery round on the discovery and held-out problems, respectively, for PersonTTS and the variant without cross-user reuse on AIME and HMMT. Held-out evaluations are used only for analysis and never exposed to the discovery agent or used for policy selection.
Bank Size
AIME24–25
AIME26
HMMT24
HMMT25
(Discovery)
(Held-out)
(Discovery)
(Held-out)
20
92.24
80.30
86.22
86.09
60
94.51
81.38
87.40
85.91
100
96.57
83.85
90.86
77.28
Table 3: Experience-bank scaling: JSR (%; higher is better) of the final discovery-selected policies. Scores are averaged over available bank-seed runs within each target profile, then equally over 20 profiles. Bold marks the highest reported value in each column. The 100-profile bank is the full source bank used in the main comparison.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Target
Baseline mean
Baseline max
Matched transfer
PersonTTS
HMMT25 ( ← AIME)
4.14
11.51
25.10
77.28
AIME26 ( ← HMMT)
10.14
35.47
43.04
83.85
Appendix
Table D.1: Cross-benchmark evaluation: JSR (%; higher is better), averaged over 20 target profiles with 1,000 replay seeds each. Baseline mean/max summarize the five external configurations in Table 1 plus AutoTTS-template, which is not reported there. PersonTTS reports the target-native result on the corresponding target profiles and held-out problems in that table.
Seed
π : (Auπ,Luπ,Cuπ)
π′ : (Auπ′,Luπ′,Cuπ′)
1
(1,2,1)
(1,1,1)
2
(1,3,1)
(1,3,3)
3
(0.5,1,3)
(0.5,2,1)
Appendix
Table E.1: Synthetic source records with identical marginal distributions. Each triple is (batch accuracy, maximum question latency, mean question cost).
Existing approaches to LLM personalization focus on constructing better personalized models or inputs, while treating inference as a single-shot process. In this work, we study Test-Time Personalization (TTP) along an unexplored axis: scaling inference-time computation by sampling N candidates from a personalized policy model and selecting the best with a personalized reward model. We prove that oracle selection yields expected utility growing logarithmically with the number of sampled candidates, establishing a theoretical ceiling for test-time scaling. However, standard reward models fail to realize this potential. To diagnose why, we derive a unified scaling law that decomposes any reward model's Best-of-N curve into four measurable quantities and reveals two failure modes, user-level collapse (near-constant prediction for some users) and query-level reward hacking (negative correlation with true quality for some queries). Guided by this law, we propose a probabilistic personalized reward model whose learned variance effectively mitigates both failure modes. Experiments confirm both elements of our framework: TTP delivers consistent scaling across multiple policy models and personalized text generation tasks, and our scaling law closely matches observed scaling curves across reward-model variants.
Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during inference. However, existing TTS strategies are largely hand-crafted: researchers manually design reasoning patterns and tune heuristics by intuition, leaving much of the computation-allocation space unexplored. We propose an environment-driven framework, AutoTTS, that changes what researchers design: from individual TTS heuristics to environments where TTS strategies can be discovered automatically. The key to AutoTTS lies in environment construction: the discovery environment must make the control space tractable and provide cheap, frequent feedback for TTS search. As a concrete instantiation, we formulate width--depth TTS as controller synthesis over pre-collected reasoning trajectories and probe signals, where controllers decide when to branch, continue, probe, prune, or stop and can be evaluated cheaply without repeated LLM calls. We further introduce beta parameterization to make the search tractable and fine-grained execution trace feedback to improve discovery efficiency by helping the agent diagnose why a TTS program fails. Experiments on mathematical reasoning benchmarks show that the discovered strategies improve the overall accuracy--cost tradeoff over strong manually designed baselines. The discovered strategies generalize to held-out benchmarks and model scales, while the entire discovery costs only $39.9 and 160 minutes. Our data, and code will be open-source at https://github.com/zhengkid/AutoTTS.
Test-Time Scaling (TTS) enhances the reasoning capabilities of large language models by allocating additional inference compute to explore the solution space. However, existing parallel TTS methods typically keep branches isolated during search: intermediate discoveries remain branch-private and cannot guide other branches in time. This information isolation causes substantial redundant exploration, as branches repeatedly rediscover information already found elsewhere and require more search steps to collect complete decision information needed to reach correct answers. To bridge this gap, we propose \textbf{Collaborative Parallel Thinking (CPT)}, a training-free inference framework that enables search-time information sharing across parallel branches. CPT extracts compact intermediate information from ongoing branches, maintains a deduplicated query-level information pool, and broadcasts pool entries through the input context, allowing each branch in subsequent search steps to reuse discoveries made by other branches rather than rediscover the same information. Empirically, experiments on HMMT and AIME benchmarks show that CPT establishes a stronger accuracy--latency Pareto frontier than strong baselines across rollout budgets and model scales, highlighting search-time collaboration as an effective direction for efficient parallel TTS.
Xinglin Wang, Hao Lin, Shaoxiong Feng +9
School of Computer Science, Beijing Institute of Technology · Xiaohongshu Inc