Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.
Figures & tables
Figure 1: Comparison of Existing Test-Time Scaling Paradigms and Personalized Test-Time Scaling. Existing methods primarily optimize accuracy against latency or inference cost separately, whereas Personalized TTS targets their joint satisfaction under user-specific requirements. PersonTTS amortizes user-conditioned controller discovery by reusing requirement-matched controllers and procedural guidance from prior searches.
Method
AIME24-25
AIME26
HMMT24
HMMT25
(Discovery)
(Held-out)
(Discovery)
(Held-out)
Source
Target
Source
Target
Source
Target
Source
Target
ASC
13.53
0.14
7.03
0.00
11.91
4.89
12.03
4.98
ESC
13.28
0.14
6.22
0.00
11.22
4.89
11.83
4.98
ParallelProbeSR
23.77
10.19
23.89
9.35
23.81
19.30
18.07
11.51
AutoTTS ( β=0.5 )
12.11
16.41
10.52
16.03
7.04
2.96
7.25
3.23
Table 1: JSR (%; higher is better) across problem and profile splits. Source and target profiles follow the Setup, and problem splits are labeled separately. PersonTTS variants use the policy with the highest JSR on discovery problems across initialization and all five rounds. PersonTTS uses both reuse mechanisms. “w/o reuse” removes both the Guide and warm-start, while the other “w/o” variants remove the named mechanism. Dashes denote unreported results. Bold marks the highest reported value in each column.
Figure 2: Best-observed JSR on discovery sets across rounds. Curves show the cumulative maximum JSR on the discovery sets after initialization and each of the five candidate rounds (R0–R4), averaged over target profiles on AIME24–25 and HMMT24. The four curves compare PersonTTS with the variants without warm-start, without the Guide, and without both mechanisms.
Method
Round
Total
R0
R1
R2
R3
R4
AIME
PersonTTS (w/o reuse)
18.67 ∣ 1.58
17.37 ∣ 2.29
14.56 ∣ 2.20
12.33 ∣ 1.94
11.70 ∣ 2.29
74.63 ∣ 10.30
PersonTTS (w/o Guide)
12.74 ∣ 1.61
13.72 ∣ 2.03
14.70 ∣ 2.36
12.85 ∣ 2.27
12.31 ∣ 2.37
66.33 ∣ 10.64
PersonTTS (w/o warm-start)
13.60 ∣ 1.24
7.05 ∣ 1.29
7.56 ∣ 1.46
6.13 ∣ 1.28
5.35 ∣ 1.37
39.69 ∣ 6.64
PersonTTS
8.49 ∣ 1.15
8.90 ∣ 1.51
8.49 ∣ 1.67
7.83 ∣ 1.57
6.70 ∣ 1.50
40.41 ∣ 7.40
Table 2: Agent-call elapsed time and cost during target-profile discovery on discovery problems. Each cell reports time (minutes) ∣ cost (USD), with lower values better for both metrics. Values are averaged over 20 target profiles per benchmark and round. Total gives the per-profile sum across the five rounds. Variant names follow Table 1 .
Figure 3: Discovery-to-held-out generalization across rounds. Solid and dashed lines report the JSR of the policy produced at each discovery round on the discovery and held-out problems, respectively, for PersonTTS and the variant without cross-user reuse on AIME and HMMT. Held-out evaluations are used only for analysis and never exposed to the discovery agent or used for policy selection.
Bank Size
AIME24–25
AIME26
HMMT24
HMMT25
(Discovery)
(Held-out)
(Discovery)
(Held-out)
20
92.24
80.30
86.22
86.09
60
94.51
81.38
87.40
85.91
100
96.57
83.85
90.86
77.28
Table 3: Experience-bank scaling: JSR (%; higher is better) of the final discovery-selected policies. Scores are averaged over available bank-seed runs within each target profile, then equally over 20 profiles. Bold marks the highest reported value in each column. The 100-profile bank is the full source bank used in the main comparison.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Target
Baseline mean
Baseline max
Matched transfer
PersonTTS
HMMT25 ( ← AIME)
4.14
11.51
25.10
77.28
AIME26 ( ← HMMT)
10.14
35.47
43.04
83.85
Appendix
Table D.1: Cross-benchmark evaluation: JSR (%; higher is better), averaged over 20 target profiles with 1,000 replay seeds each. Baseline mean/max summarize the five external configurations in Table 1 plus AutoTTS-template, which is not reported there. PersonTTS reports the target-native result on the corresponding target profiles and held-out problems in that table.
Seed
π : (Auπ,Luπ,Cuπ)
π′ : (Auπ′,Luπ′,Cuπ′)
1
(1,2,1)
(1,1,1)
2
(1,3,1)
(1,3,3)
3
(0.5,1,3)
(0.5,2,1)
Appendix
Table E.1: Synthetic source records with identical marginal distributions. Each triple is (batch accuracy, maximum question latency, mean question cost).