How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@K tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end. Single-shot or short-horizon tasks avoid these tool-calling failures by collapsing a multi-step interaction into a fixed prompt and a single patch, but they sidestep the core capability we care about: maintaining coherent state over many tool-using steps as the repository evolves. To bridge this gap, we treat successful post-trained agent trajectories as a lookahead signal of base-model potential. Replaying each trajectory and rerunning tests after every code-changing step identifies the decisive step: the first step whose cumulative patch flips the repository from failing to passing, certifying that the recorded action solves the task given the prior context. Motivated by a coverage principle for agentic traces, we build three screens at this step that do not require a base checkpoint to drive the harness from a cold start: (i) Decisive-Action BPB (bits per byte) measures the probability mass on the certified action, (ii) Patch MCQ tests the checkpoint's choice between that action and alternatives rejected by the same verifier, and (iii) prefix-conditioned pass@K evaluates support for functionally-correct generations and credits any continuation that the tests accept. Across ten pairs of public base and post-trained models, all three screens rank the cohort in close agreement with post-trained SWE-bench Verified pass@1. As our methods need only a benchmark's successful trajectories and its verifier, they can be applied to turn future agentic coding benchmarks into base-model evaluations.
Figures & tables
Figure 1: A unified agentic trajectory-derived probing framework. By replaying a coding-agent trajectory on a coding task against the task’s own verifier V , we locate the decisive step T⋆ , the earliest step whose cumulative patch resolves the task. At this decisive step, (i) Decisive-Action BPB measures the byte-normalized likelihood which a base model assigns to the golden action; (ii) Patch MCQ evaluates base model’s capability of differentiating the golden action from non-resolving alternatives; (iii) Prefix-conditioned pass@K measures a base model’s empirical success rate under sampling: it draws K continuations from T⋆ and lets the task’s own verifier V judge them.
Figure 2: (a) Restricting BPB to the verifier-derived decisive step, instead of averaging over every step, raises its Spearman’s rank correlation ( ρ ) with downstream agent capability, from 0.915 to 0.964 against SWE-bench Verified, across the same ten checkpoints. (b) At the checkpoint level, Decisive-Action BPB ranks DeepSeek V4 Flash above Nemotron Ultra and Kimi K2 above Qwen 3.5 35B, in agreement with their downstream performance, thus reversing the incorrect orderings produced by all-step BPB.
Figure 3: All three DeepSWE-derived probes strongly track post-trained SWE-bench Verified performance. The common DeepSWE source corpus makes the panels directly comparable across Decisive-Action BPB with tool-formatting tokens masked, Patch MCQ, and prefix-conditioned pass@K at K=32 . The same ten checkpoints appear in every panel; annotations report Pearson r and Spearman ρ . The BPB axis is reversed so rightward consistently means better.
Figure 6: The three DeepSWE-derived probes retain strong rank agreement with post-trained SWE-bench Multilingual performance. Panels use the same ten checkpoints and the same probe settings as fig. 5 ; annotations report Pearson r and Spearman ρ . The BPB axis is reversed so rightward consistently means better.
Post-SFT
Public
Decisive-Action
Patch
Prefix
Base checkpoint
SWE-Verified
SWE-Verified
BPB
MCQ
pass@32
Nemotron-3-Nano-Base
28.5
38.8
0.149
33.4
13.1
Nemotron-3-Super-Base
53.1
60.5
0.100
34.4
34.6
Qwen-3.5-35B-A3B-Base
61.2
69.2
0.092
36.2
46.9
Table 1: Under the same SFT procedure, post-SFT SWE-bench Verified follows the same ordering as the publicly reported scores, and the probes order the checkpoints consistently with it. Probe results are computed with SWE-Pro trajectories.
Figure 7: Masking tool-call formatting and scoring only decisive actions improve the BPB screen. Bars report Spearman correlation with post-trained SWE-bench Verified; all correlations use −BPB so higher is better.
Table 7
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Component
All questions
Invalid questions
Invalid (%)
Fix Patch
232
63
27.2%
Plan
243
81
33.3%
Test Patch
1,060
450
42.5%
Handle Error
147
133
90.5%
Total
1,682
727
43.2%
Appendix
Table 4: Rubric-based assessment of 1,682 APTBench SWE MCQs. An MCQ is classified as Invalid when the question does not have a uniquely correct designated answer and incorrect distractors. Percentages are relative to all questions in each row.
SWE-bench Verified
APTBench metric
Pearson
Spearman
APTBench-SWE average
0.942
0.879
APTBench-DR average
0.962
0.903
APTBench-SWE: EnvSetup Plan ACC
0.910
0.818
APTBench-SWE: EnvSetup Action EM
0.851
0.842
APTBench-SWE: EnvSetup Handle Error ACC
0.871
0.770
Appendix
Table 5: APTBench reproduction on the same ten base checkpoints evaluated by the three probes ( section 4.1 ). Pearson and Spearman correlations between APTBench scores and SWE-bench Verified performance of the corresponding post-trained release for each checkpoint.
SWE-bench
SWE-bench
Terminal-Bench
Base checkpoint
Post-trained model
Verified
Multilingual
2.1
Nemotron-Nano-Base
Nemotron-Nano
38.8
33.3
6.7
Nemotron-Super-Base
Nemotron-Super
60.5
45.8
38.6
Nemotron-Ultra-Base
Nemotron-Ultra
70.7
67.7
53.9
Qwen-3.5-35B-A3B-Base
Qwen-3.5-35B-A3B
69.2
60.3
48.3
DeepSeek-V4-Flash-Base
DeepSeek-V4-Flash
79.0
73.3
78.7
Appendix
Table 6: Evaluation cohort and downstream targets: SWE-bench Verified ( Jimenez et al., 2024 ; OpenAI, 2024 ) and Terminal-Bench 2.1 ( Merrill et al., 2026 ) . Each row pairs a base checkpoint with the post-trained release whose reported performance we use as the downstream target; naming it matters because several families shipped more than one. All targets are pass@1 percentages.
Figure 8: The pipeline shared by all three probes. Stages 1–5 run once per trajectory corpus and produce artifacts reused for every checkpoint; stage 6 is the only per-checkpoint work, and it is where the static and sampled probes diverge by orders of magnitude. Verifier executions, not forward passes, are the scarce resource, which is why stage 4 bisects rather than scans and why the sampled path caches prefixes.
vs. SWE-bench Verified
vs. Terminal-Bench 2.1
Evaluation
n
Pearson
Spearman
Pearson
Spearman
MBPP
10
0.151
0.091
0.037
0.036
HumanEval
10
−0.369
−0.394
−0.650
−0.571
LiveCodeBench
10
0.523
0.455
0.429
0.383
CRUXEval-I
10
0.473
0.394
0.287
0.201
CRUXEval-O
10
0.814
0.758
0.618
0.523
Appendix
Table 7: Bounded code benchmarks as baselines for post-trained agentic performance, with our three probes repeated for reference. Correlations are against each post-trained downstream target on the cohort of table 6 ; the probe rows use the DeepSWE-derived probes of figs. 5 and 9 , with −BPB for Decisive-Action BPB. Two baselines cover only ten checkpoints, so their coefficients are not matched pairs with the rest. At n≤10 with no interval estimates these are descriptive associations, and single-row differences should not be over-read.
End-to-end
Post-trained
Base checkpoint
pass@1
SWE-Verified
DeepSeek-V4-Pro-Base
0.0
80.6
DeepSeek-V4-Flash-Base
0.0
79.0
Nemotron-Ultra-Base
0.0
70.7
Qwen-3.5-35B-A3B-Base
20.2
69.2
Nemotron-Super-Base
0.0
60.5
Appendix
Table 8: End-to-end evaluation of base checkpoints on SWE-bench Verified, alongside each family’s post-trained score on the same benchmark. All values are pass@1 percentages.
Figure 9: DeepSWE prefix-conditioned pass@K at K=32 has the strongest descriptive agreement with post-trained Terminal-Bench 2.1 performance. Panels use the same ten checkpoints and the same probe settings as fig. 5 ; annotations report Pearson r and Spearman ρ . The BPB axis is reversed so rightward consistently means better.
Source component
Universal representation
Tool definitions
Tool definitions: <JSON>
System message
System: <content>
User message
User: <content>
Assistant text
Assistant: <content>
Assistant tool call
Assistant (tool call): <JSON>
Tool response
Tool Result: <content>
Appendix
Table 9: Universal serialization of trajectory components. Angle-bracketed text denotes the original message content.
DeepSWE
SWE-Verified
Base checkpoint
No think
Think
No think
Think
Hy3-Preview-Base
0.4456
0.5080
0.552
0.583
DeepSeek-V4-Pro-Base
0.4206
0.5026
0.550
0.591
Nemotron-Ultra-Base
0.4235
0.4611
0.538
0.558
DeepSeek-V4-Flash-Base
0.4009
0.4421
0.536
0.553
Kimi-K2-Base
0.4322
0.4047
0.597
0.523
Appendix
Table 10: Source-current four-option Patch MCQ accuracy with cyclic rotation and arithmetic aggregation; chance is 0.250. Both question sets use matched none preambles. Each cell is the one accuracy our source measurements record for that checkpoint and setting; run-to-run spread is not stored with these values ( appendix C ), so column-to-column gaps of a few thousandths should not be read as differences. Bold marks the largest value per column.
DeepSWE
SWE-Verified
Base checkpoint
Argmin
Rotated
Gap
Argmin
Rotated
Gap
Hy3-Preview-Base
0.266
0.508
−0.242
0.359
0.583
−0.224
DeepSeek-V4-Pro-Base
0.283
0.503
−0.220
0.405
0.591
−0.186
Nemotron-Ultra-Base
0.267
0.461
−0.194
0.366
0.558
−0.192
DeepSeek-V4-Flash-Base
0.266
0.442
−0.176
0.380
0.553
−0.173
Kimi-K2-Base
0.279
0.405
−0.126
0.387
0.523
−0.136
Appendix
Table 11: Patch MCQ read ablation on both question sets; chance is 0.250. Argmin-BPB scores each option independently and picks the lowest byte-normalized negative log-likelihood; rotated log-probability is the answer-letter read, presenting all four options jointly under four cyclic rotations, and is reported here at think-1000 to match table 10 . Gap is argmin-BPB minus rotated log-probability, so negative values favor the joint read. Rows are ordered by DeepSWE rotated log-probability.
Base checkpoint
k=1
k=4
Gain
DeepSeek-V4-Pro-Base
0.4555
0.5026
+0.047
Hy3-Preview-Base
0.4271
0.5080
+0.081
Nemotron-Ultra-Base
0.4195
0.4611
+0.042
DeepSeek-V4-Flash-Base
0.3881
0.4421
+0.054
Kimi-K2-Base
0.3741
0.4047
+0.031
Nemotron-Super-Base
0.3727
0.4116
+0.039
Appendix
Table 12: Single-rotation versus full cyclic rotation on the DeepSWE question set, think-1000 arm, over the same ten checkpoints as table 10 ; chance is 0.250. k=1 takes the answer from the first rotation alone and k=4 is the shipped aggregate, both recomputed offline from the same finished runs, so no checkpoint was re-scored; each value is the mean of three rounds. The k=4 column is the DeepSWE think-1000 column of table 10 . Rotation helps on 9 of 10 checkpoints, but not uniformly: the gain spans −0.005 to +0.081 , wider than the gap between adjacent checkpoints in the k=4 ranking, and it reorders the top two. Bold marks the best value per column.
vs. SWE-bench Verified
vs. Terminal-Bench 2.1
Question set
Scoring method
Pearson
Spearman
Pearson
Spearman
DeepSWE
Argmin-BPB
0.601
0.665
0.420
0.382
DeepSWE
Rotated, no think
0.900
0.806
0.686
0.547
DeepSWE
Rotated, think 1000
0.915
0.867
0.871
0.796
SWE-Verified
Argmin-BPB
0.429
0.578
0.359
0.213
SWE-Verified
Rotated, no think
0.909
0.842
0.725
0.578
Appendix
Table 13: Downstream agreement of every Patch MCQ scoring method on the same ten checkpoints, with both Pearson and Spearman correlations. Rows pair a question set with a scoring method; the SWE-Verified question set draws its options from the same benchmark as the first target, so those rows are same-benchmark evidence against it while all rows are cross-domain with respect to Terminal-Bench 2.1. At n=10 with no interval estimates we do not read small differences as separating rows.
Figure 10: Patch MCQ accuracy and downstream agreement, by scoring rule. Panels (a) and (b) rank the same ten checkpoints on each question set; panels (c) and (d) plot the rank agreement tabulated in table 13 against both downstream targets. The joint answer-letter read separates the cohort more sharply than scoring each option independently, and the gap between the two reads is larger than the gap between thinking protocols.
Figure 11: Driving an agent harness from a base checkpoint. (a) The bridge sits between an unmodified agent CLI and the base checkpoint: it renders each agent request into a plain-text transcript, samples a continuation from the checkpoint’s raw completion endpoint, and returns the parsed action to the agent as a native tool call. (b) The prompt at one forward step, assembled as transcript, steer, primer, and prefill. The prefill pre-opens the action envelope so that the checkpoint’s first sampled token already lands inside an open action; only the green span is written by the base checkpoint.
DeepSWE
SWE-bench Verified
Base checkpoint
k=8
k=16
k=32
k=8
k=16
k=32
hy3-base
16.67
20.83
26.67
61.22
70.41
76.53
deepseek-v4-pro-base
37.50
41.67
45.83
68.37
77.55
83.67
ultra-base
12.50
17.50
26.67
58.16
68.37
83.67
dsflash-base
28.33
33.33
40.00
63.27
72.45
79.59
kimi-k2-base
14.17
16.67
21.67
55.10
62.24
68.37
Appendix
Table 14: Prefix-conditioned pass@K on DeepSWE and SWE-bench Verified trajectory prefixes, in percent. DeepSWE uses 120 held-out instances; SWE-bench Verified uses a separately eligibility-filtered cohort of 98 prefixes. Both cover the same ten checkpoints, but the two halves aggregate over different cohort sizes, so their means are not matched pairs. Rollouts fork at the decisive step Ti⋆ : steps before are served verbatim, Ti⋆ is masked, and the checkpoint keeps control for the rest of the episode. An instance counts as resolved if at least one of k continuations passes the verifier. Row order follows table 10 ; bold marks the largest value per column.
vs. SWE-bench Verified
vs. Terminal-Bench 2.1
Probe
n
Pearson
Spearman
Pearson
Spearman
DeepSWE prefix pass@K , k=8
10
0.840
0.915
0.870
0.802
DeepSWE prefix pass@K , k=16
10
0.872
0.952
0.912
0.900
DeepSWE prefix pass@K , k=32
10
0.917
0.951
0.947
0.930
SWE-bench Verified prefix pass@K , k=8
10
0.990
0.988
0.902
0.875
SWE-bench Verified prefix pass@K , k=16
10
0.984
0.988
0.911
0.875
Appendix
Table 15: Downstream validation of prefix-conditioned pass@K against both targets of table 6 , on the same ten-checkpoint cohort used throughout the primary comparisons; higher is better agreement. SWE-bench Verified prefix rows are same-benchmark evidence against the SWE-bench Verified target, not cross-domain validation; both prefix corpora are cross-domain with respect to Terminal-Bench 2.1. The last two rows repeat the canonical all-step/decisive-step comparison from figs. 2 and 7 , correlated against −BPB on the same ten checkpoints. No value is marked best: at n=10 with no bootstrap intervals, we do not read these differences as separating the rows.
Figure 12: Prefix-conditioned pass@K rankings and downstream agreement. Panels (a) and (b) rank the same ten checkpoints under DeepSWE and SWE-bench Verified trajectory prefixes at k∈{8,16,32} , tabulated in table 14 . The two panels are sorted independently and use different x ranges, so marker positions are not comparable across them. Panels (c) and (d) plot the rank agreement of those scores with each downstream target from table 15 ; the dashed line marks Decisive-Action BPB against the same target for reference.
vs. Post-SFT SWE-Verified
vs. Public SWE-Verified
vs. Public SWE-Verified
( n=3 )
( n=3 )
( n=10 )
Evaluation
Pearson
Spearman
Pearson
Spearman
Pearson
Spearman
DeepSWE
Decisive-Action BPB
0.97
0.50
0.96
0.50
0.97
0.96
Patch MCQ
0.91
0.50
0.90
0.50
0.87
0.80
Prefix pass@32
0.97
1.00
0.98
1.00
0.95
0.93
Appendix
Table 16: Comparison of probe correlations across source tasks and post-SFT SWE-bench Verified results for Nemotron 3 Nano, Nemotron 3 Super, and Qwen 3.5 35B A3B. We show correlations against public SWE-bench Verified scores for the same 3 models, as well as all 10 models in our panel.