How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-POT separates candidate coverage from final accuracy on static tasks, tracks correctness transitions under revision, and measures protocol completion alongside task success in agentic environments. Under a unified budget rule, it compares Direct inference with parallel sampling and self-revision under fixed multiples of the Direct budget, and charges every model call, including selection and critique, in dollars. This design supports two kinds of comparison: the gain a model obtains from additional inference, and a lower-cost model with additional inference against a stronger model. Across five low-cost reasoning models on 350 sealed tasks, with Claude Opus 5.5 Direct as the reference, the returns depend on the domain, the selection rule, and failure handling. When we replay the retained programming candidate pools, public-example selection raises correct submissions from 376 to 453 of 500 scheduled cells while saving 12-49% of logical API cost across models, and simply retaining an available candidate when judging fails recovers 61 submissions at unchanged cost. On identical mathematics pools, judging with fallback yields 186 correct submissions versus 182 for voting, while voting saves 12-21% of logical API cost. These controlled replays show how selection and failure handling change the gains realized from the same generated candidates, and they quantify the marginal value of a model judge.
Figures & tables
Figure 1 : SELF-POT : frozen protocol matrix. Each quarter (equal angles) is one domain of the 350-task sealed test set, with its sampling strata in the inner ring. Bars show each low-cost model’s Direct accuracy (%), markers the same model under the TTS protocols, and arcs Claude Opus 5.5 answering once. A marker beyond the end of a bar is a gain, and a marker inside a bar is a loss.
Figure 2 : SELF-POT protocols. One frozen low-cost model answers the same task under four protocols whose total output is capped at a multiple of the budget B of one answer, and every call is charged. On agent tasks, the parallel protocols compare plans that are executed once, and Revise@4 reviews the trajectory every two steps (§ 2.2 ). The mathematics problem is illustrative.
Domain
Source
Tasks
Feedback while solving
Actions reversible
Math
APEX [ 4 , 37 ] , Omni-MATH [ 20 ]
60
none
yes
Code
LiveCodeBench v6 [ 27 ] (AtCoder)
100
public examples and execution
yes
Workflow
AppWorld [ 50 ]
160
app API responses
no
Terminal
Terminal-Bench 4.0 [ 38 ]
30
shell output
no
Table 1: The four domains of the sealed test set. Hidden answers, tests, and goal checks are used only for grading; Appendix B gives the sampling strata.
Model
Access
Price (in / out)
Context
Protocols
DeepSeek V4.1 Flash
DeepSeek API
0.30 / 1.20 a
1M
all
Qwen3-Next-80B-A3B
Amazon Bedrock
0.15 / 1.20
262,144
all
GLM-5
Amazon Bedrock
1.00 / 3.20
202,752
Direct, Parallel
GPT-OSS-120B
Amazon Bedrock
0.1545 / 0.618
131,072
Direct, Parallel b
MiniMax M2.5
Amazon Bedrock
0.30 / 1.20
196,608
Direct, Parallel b
Claude Opus 5.5 (anchor)
Amazon Bedrock
4.00 / 20.00
–
Direct
Table 2: Models. Prices are USD per million input and output tokens on our ledger. Revision eligibility uses the configured input headroom plus a full-budget critique, not a universal context threshold for all revision methods.
Direct
Parallel@2
Parallel@4
Revise@4
Domain
Model
Acc
$
Acc
$
Acc
$
Acc
$
Math (60)
Opus 5.5
0.950
0.103
–
–
–
–
–
–
DeepSeek V4.1 Flash
0.767
0.051
0.783
0.098
0.867
0.176
0.800
0.096
GPT-OSS-120B
0.533
0.026
0.533
0.052
0.583 ‡
0.096
–
–
MiniMax M2.5
0.417
0.038
0.417
0.078
0.450
0.151
–
–
GLM-5
0.650
0.175
0.667 ‡
0.347
0.683 ‡
0.618
–
–
Table 3: Frozen protocol matrix: accuracy (Acc) over all scheduled tasks and mean logical API cost in USD over retained cell records. Missing records have no imputed cost. Opus 5.5 is the expensive anchor and runs Direct only. Bold marks the best accuracy among the four displayed protocols in each row; “–” means not run; ‡ includes cells that never finished, counted as incorrect.
Figure 3 : Frozen protocol matrix: accuracy change against the same model’s Direct run, in percentage points (large) and paired wins–losses (small). Stars: uncorrected exact sign tests ( ∗p<0.05 , ∗∗p<0.01 ); Holm-adjusted tests are in Table 9 . n/a: excluded from revision by context; —: not run.
Parallel@2
Parallel@4
Revise@4
Domain
Δ
95% interval
W–L
Δ
95% interval
W–L
Δ
95% interval
W–L
Math
+0.7
[−0.7,+2.0]
3–1
+4.0
[+0.3,+7.7]
24–12
+0.8
[−1.7,+3.3]
2–1
Code
+6.2
[+4.2,+8.2]
31–0
−8.6
[−11.8,−5.4]
18–61
−0.5
[−1.5,+0.0]
0–1
Workflow
−4.2
[−7.2,−1.4]
49–83
−4.6
[−7.4,−1.9]
48–85
−7.2
[−11.2,−3.4]
11–34
Terminal
+0.0
[−5.6,+5.6]
3–3
+0.0
[−4.4,+4.4]
2–2
−5.0
[−10.0,+0.0]
0–3
Table 4: Frozen protocol matrix: mean paired accuracy change Δ (pp) against Direct over the fixed panel of low-cost models (five for Parallel, three on terminal tasks, two for Revise@4), with pointwise 95% task-bootstrap intervals and pooled wins–losses (W–L). The intervals describe this panel, not unseen model families.
Figure 4 : Candidate coverage and conversion under frozen Parallel@4. Each bar partitions scheduled tasks. White numbers count correct submissions, labels above bars count covered tasks, and black ticks mark Direct accuracy. “No choice” denotes no valid index; “wrong choice” includes selecting an empty candidate. Same-pool alternatives are compared in Figure 5 and Appendix J.2 .
Figure 5 : Same-pool choices and logical API cost. Accuracy uses scheduled denominators (60 mathematics or 100 coding tasks per model), retaining missing records as failures; cost averages over retained records. Mathematics voting and programming public selection remove judge-stage calls. Fallback retains already-spent cost. Gate costs assume calling the unchanged judge only on public-score ties (Appendix J.2 ). Local execution expense is reported separately.
Figure 6 : Cost and accuracy of the frozen configurations; rings mark their Pareto frontier. Selector replays are compared separately in Figure 5 . Costs are mean logical API costs per task over retained cells (DeepSeek peak rates; Bedrock list rates, with Bedrock caching only for Opus agent requests). Code uses the official grades (Opus 93%, 99% with numpy ), and very low workflow costs can reflect early protocol failures. The frontiers are point estimates, not certified dominance relations.
Study
Domains
Methods
Judge charged
Budget unit
Reasoning models
vs. stronger
Snell et al. [47]
M
P, S
✗
samples, FLOPs
✗
✓
Wu et al. [53]
M, C
P
✗
FLOPs
✗
✓
Wang et al. [51]
M, O
P, S
✓
tokens, $
✗
✗
Chen et al. [7]
O
P
✗
calls
✗
✗
Singhi et al. [46]
M, O
P
✓
FLOPs
✓
✗
Ghosal et al. [21]
M
P, S
—
tokens
✓
✗
Table 5: Controlled studies of test-time scaling. M/C/A/O: mathematics, code, agent tasks, other; P/S: parallel, sequential. Judge charged : selection calls count against the generation budget (—: no judge call). vs. stronger : a cheaper model with more compute is compared with a stronger model.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Protocol
Generation
Selection or feedback
Total cap
Direct
B
–
B
Parallel@2
2×B
rule, no model tokens
2B
Parallel@4
3×B
joint selection, B
4B
Revise@4
initial B , revision up to 2B
critique, B
4B
Appendix
Table 6: Nominal stage allowances for static tasks. The suffix denotes a total budget multiplier, not a sample count. Provider limits can reduce individual allowances; in particular, Qwen3-Next’s revision is capped at B , giving at most 3B across its three stages.
Domain
Output tokens B
Model calls
Tool calls
Wall clock
Math
131,072
16
–
2 h
Code
131,072
64
128
2 h
Workflow
393,216 a
48
64
1 h
Terminal
393,216
128
256
8 h b
Appendix
Table 7: Per-cell caps for Direct ( m=1 ). All caps are multiplied by m for Parallel@2 ( m=2 ), Parallel@4 and Revise@4 ( m=4 ).
Table 9: Per-model comparisons with Holm-adjusted p<0.05 across the full family of 44 protocol-versus-Direct tests. The family includes every evaluated non-Direct configuration, not only displayed effects.
Model
Domain
R → R
W → R
R → W
W → W
Recovery
Loss
DeepSeek V4.1 Flash
Math
46
2
0
12
2/14
0/46
Qwen3-Next-80B
Math
30
0
1
29
0/29
1/31
DeepSeek V4.1 Flash
Code
98
0
1
1
0/1
1/99
Qwen3-Next-80B
Code
76
0
0
24
0/24
0/76
Appendix
Table 10: Static Revise@4 transitions from its own recorded initial answer to its final submission. R and W denote correct and incorrect, including missing or ungradable outputs. Recovery and loss are operational rates, not solely content changes.
Model
Pools
Coverage
Self-select
Vote@3
Conversion
DeepSeek V4.1 Flash
60
54
52
50
52/54
Qwen3-Next-80B
60
38
30
31
30/38
GLM-5
53
44
41
41
41/44
GPT-OSS-120B
59
42
35
34
35/42
MiniMax M2.5
60
31
27
26
27/31
Appendix
Table 11: Mathematics Parallel@4: coverage, self-selection, and counterfactual vote@3 on the identical recorded pools. Every available pool has three graded candidates. Counts in the middle columns use the available-pool denominator; official accuracy always uses all 60 scheduled tasks.
Model
3 attempts
O
S
No choice
Wrong choice
DeepSeek V4.1 Flash
100
99
96
3
0
Qwen3-Next-80B
100
83
75
1
7
GLM-5
100
88
55
27
6
GPT-OSS-120B
96
94
69
25
0
MiniMax M2.5
100
93
81
8
4
Appendix
Table 12: Programming Parallel@4: where recorded candidate coverage is lost. Each model has 100 scheduled tasks. O : a correct candidate is observed; S : the final submission is correct. The last two columns partition O−S . A missing record is not counted as an observed empty pool.
Model
Direct
J
J+F
J+P
P3
G
P2
Flash
99
96
99
99
99
99
99
Qwen3-Next
76
75
76
76
82
81
82
GLM-5
66
55
78
81
87
87
83
GPT-OSS
91
69
94
94
94
94
94
MiniMax
87
81
90
90
91
91
90
Total
419
376
437
440
453
452
448
Appendix
Table 13: Programming selector replay on the recorded Parallel@4 pools. Counts use 100 scheduled tasks per model. J: frozen judge; J+F: first-nonempty fallback; J+P: public-example fallback (both also handle an empty chosen answer); P3/P2: public-example choice on all three/first two candidates; G: public tests first, judge only tied maxima, with public fallback.
Model
J / J+F / J+P
P3
P2
G
Flash
0.0766
0.0391
0.0269
0.0766
Qwen3-Next
0.0800
0.0566
0.0381
0.0792
GLM-5
0.3774
0.3058
0.2096
0.3647
GPT-OSS
0.0809
0.0709
0.0533
0.0793
MiniMax
0.1686
0.1172
0.0750
0.1646
Appendix
Table 14: Logical API costs of the coding replay (USD per retained task). Fallback retains all judge spending; public-only variants omit it, and G charges the judge only for public-score ties.
Candidate pool
Pools
J
J+F
P3
All correct
359
315
359
359
Mixed
98
61
78
94
All incorrect
41
0
0
0
Appendix
Table 15: Programming conversion by recorded candidate correctness. J: frozen judge; J+F: nonempty fallback; P3: public-example selection. Counts exclude the two missing records, which remain failures in scheduled accuracy.
Model
Pools
J
J+F
Vote
J / J+F $
Vote $
Saving
Flash
60
52
52
50
0.1764
0.1444
18.1%
Qwen3-Next
60
30
30
31
0.1561
0.1369
12.3%
GLM-5
53
41
41
41
0.6183
0.4882
21.1%
GPT-OSS
59
35
35
34
0.0959
0.0758
21.0%
MiniMax
60
27
28
26
0.1511
0.1193
21.0%
Appendix
Table 16: Mathematics original judging (J), judging with nonempty fallback (J+F), and vote@3 on identical pools. Costs are mean logical API dollars per retained task at the original price basis; voting removes the entire judge stage.
Grader
Pools
J
J+F
V
W–L
Primary rule
292
185
186
182
10–6
Compass
292
192
194
189
11–6
Appendix
Table 17: Mathematics selector sensitivity under two grading rules on identical candidates and fixed choices. J: frozen judge; J+F: nonempty fallback; V: vote@3. W–L compares J+F with V.
Direct
Parallel@2
Parallel@4
Revise@4
Anchor
Domain
Model
Acc
$/task
Acc (W–L)
Cost
Acc (W–L)
Cost
Acc (W–L)
Cost
Opus 5.5
Math (60)
DeepSeek V4.1 Flash
0.767
0.051
0.783 (1–0)
1.9 ×
0.867 (6–0) ∗
3.4 ×
0.800 (2–0)
1.9 ×
0.950
GPT-OSS-120B
0.533
0.026
0.533 (0–0)
2.0 ×
0.583 ‡ (5–2)
3.7 ×
n/a
n/a
MiniMax M2.5
0.417
0.038
0.417 (0–0)
2.0 ×
0.450 (3–1)
3.9 ×
n/a
n/a
GLM-5
0.650
0.175
0.667 ‡ (2–1)
2.0 ×
0.683 ‡ (7–5)
3.5 ×
n/a
n/a
Qwen3-Next-80B
0.517
0.048
0.517 (0–0)
1.9 ×
0.500 (3–4)
3.3 ×
0.500 (0–1)
1.2 ×
Appendix
Table 18: Frozen protocol matrix: accuracy of each cheap model under each protocol, with wins–losses against the same model’s Direct on the same tasks (exact two-sided sign test: ∗p<0.05 , ∗∗p<0.01 , uncorrected) and cost relative to Direct. The rightmost column gives the expensive anchor, Claude Opus 5.5, answering once. n/a: excluded by the configured revision headroom requirement. –: not run. ‡ : includes cells that never finished, scored incorrect. § : the plan or plan-selection step failed its required format in at least 69% of tasks, so the row reflects format compliance rather than the value of extra compute.