How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@K tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end. Single-shot or short-horizon tasks avoid these tool-calling failures by collapsing a multi-step interaction into a fixed prompt and a single patch, but they sidestep the core capability we care about: maintaining coherent state over many tool-using steps as the repository evolves. To bridge this gap, we treat successful post-trained agent trajectories as a lookahead signal of base-model potential. Replaying each trajectory and rerunning tests after every code-changing step identifies the decisive step: the first step whose cumulative patch flips the repository from failing to passing, certifying that the recorded action solves the task given the prior context. Motivated by a coverage principle for agentic traces, we build three screens at this step that do not require a base checkpoint to drive the harness from a cold start: (i) Decisive-Action BPB (bits per byte) measures the probability mass on the certified action, (ii) Patch MCQ tests the checkpoint's choice between that action and alternatives rejected by the same verifier, and (iii) prefix-conditioned pass@K evaluates support for functionally-correct generations and credits any continuation that the tests accept. Across ten pairs of public base and post-trained models, all three screens rank the cohort in close agreement with post-trained SWE-bench Verified pass@1. As our methods need only a benchmark's successful trajectories and its verifier, they can be applied to turn future agentic coding benchmarks into base-model evaluations.
Figures & tables
Figure 1: A unified agentic trajectory-derived probing framework. By replaying a coding-agent trajectory on a coding task against the task’s own verifier V , we locate the decisive step T⋆ , the earliest step whose cumulative patch resolves the task. At this decisive step, (i) Decisive-Action BPB measures the byte-normalized likelihood which a base model assigns to the golden action; (ii) Patch MCQ evaluates base model’s capability of differentiating the golden action from non-resolving alternatives; (iii) Prefix-conditioned pass@K measures a base model’s empirical success rate under sampling: it draws K continuations from T⋆ and lets the task’s own verifier V judge them.
Figure 2: (a) Restricting BPB to the verifier-derived decisive step, instead of averaging over every step, raises its Spearman’s rank correlation ( ρ ) with downstream agent capability, from 0.915 to 0.964 against SWE-bench Verified, across the same ten checkpoints. (b) At the checkpoint level, Decisive-Action BPB ranks DeepSeek V4 Flash above Nemotron Ultra and Kimi K2 above Qwen 3.5 35B, in agreement with their downstream performance, thus reversing the incorrect orderings produced by all-step BPB.
Figure 3: All three DeepSWE-derived probes strongly track post-trained SWE-bench Verified performance. The common DeepSWE source corpus makes the panels directly comparable across Decisive-Action BPB with tool-formatting tokens masked, Patch MCQ, and prefix-conditioned pass@K at K=32 . The same ten checkpoints appear in every panel; annotations report Pearson r and Spearman ρ . The BPB axis is reversed so rightward consistently means better.
Figure 6: The three DeepSWE-derived probes retain strong rank agreement with post-trained SWE-bench Multilingual performance. Panels use the same ten checkpoints and the same probe settings as fig. 5 ; annotations report Pearson r and Spearman ρ . The BPB axis is reversed so rightward consistently means better.
Post-SFT
Public
Decisive-Action
Patch
Prefix
Base checkpoint
SWE-Verified
SWE-Verified
BPB
MCQ
pass@32
Nemotron-3-Nano-Base
28.5
38.8
0.149
33.4
13.1
Nemotron-3-Super-Base
53.1
60.5
0.100
34.4
34.6
Qwen-3.5-35B-A3B-Base
61.2
69.2
0.092
36.2
46.9
Table 1: Under the same SFT procedure, post-SFT SWE-bench Verified follows the same ordering as the publicly reported scores, and the probes order the checkpoints consistently with it. Probe results are computed with SWE-Pro trajectories.
Figure 7: Masking tool-call formatting and scoring only decisive actions improve the BPB screen. Bars report Spearman correlation with post-trained SWE-bench Verified; all correlations use −BPB so higher is better.
Table 7
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Component
All questions
Invalid questions
Invalid (%)
Fix Patch
232
63
27.2%
Plan
243
81
33.3%
Test Patch
1,060
450
42.5%
Handle Error
147
133
90.5%
Total
1,682
727
43.2%
Appendix
Table 4: Rubric-based assessment of 1,682 APTBench SWE MCQs. An MCQ is classified as Invalid when the question does not have a uniquely correct designated answer and incorrect distractors. Percentages are relative to all questions in each row.
SWE-bench Verified
APTBench metric
Pearson
Spearman
APTBench-SWE average
0.942
0.879
APTBench-DR average
0.962
0.903
APTBench-SWE: EnvSetup Plan ACC
0.910
0.818
APTBench-SWE: EnvSetup Action EM
0.851
0.842
APTBench-SWE: EnvSetup Handle Error ACC
0.871
0.770
Appendix
Table 5: APTBench reproduction on the same ten base checkpoints evaluated by the three probes ( section 4.1 ). Pearson and Spearman correlations between APTBench scores and SWE-bench Verified performance of the corresponding post-trained release for each checkpoint.
SWE-bench
SWE-bench
Terminal-Bench
Base checkpoint
Post-trained model
Verified
Multilingual
2.1
Nemotron-Nano-Base
Nemotron-Nano
38.8
33.3
6.7
Nemotron-Super-Base
Nemotron-Super
60.5
45.8
38.6
Nemotron-Ultra-Base
Nemotron-Ultra
70.7
67.7
53.9
Qwen-3.5-35B-A3B-Base
Qwen-3.5-35B-A3B
69.2
60.3
48.3
DeepSeek-V4-Flash-Base
DeepSeek-V4-Flash
79.0
73.3
78.7
Appendix
Table 6: Evaluation cohort and downstream targets: SWE-bench Verified ( Jimenez et al., 2024 ; OpenAI, 2024 ) and Terminal-Bench 2.1 ( Merrill et al., 2026 ) . Each row pairs a base checkpoint with the post-trained release whose reported performance we use as the downstream target; naming it matters because several families shipped more than one. All targets are pass@1 percentages.
Figure 8: The pipeline shared by all three probes. Stages 1–5 run once per trajectory corpus and produce artifacts reused for every checkpoint; stage 6 is the only per-checkpoint work, and it is where the static and sampled probes diverge by orders of magnitude. Verifier executions, not forward passes, are the scarce resource, which is why stage 4 bisects rather than scans and why the sampled path caches prefixes.
vs. SWE-bench Verified
vs. Terminal-Bench 2.1
Evaluation
n
Pearson
Spearman
Pearson
Spearman
MBPP
10
0.151
0.091
0.037
0.036
HumanEval
10
−0.369
−0.394
−0.650
−0.571
LiveCodeBench
10
0.523
0.455
0.429
0.383
CRUXEval-I
10
0.473
0.394
0.287
0.201
CRUXEval-O
10
0.814
0.758
0.618
0.523
Appendix
Table 7: Bounded code benchmarks as baselines for post-trained agentic performance, with our three probes repeated for reference. Correlations are against each post-trained downstream target on the cohort of table 6 ; the probe rows use the DeepSWE-derived probes of figs. 5 and 9 , with −BPB for Decisive-Action BPB. Two baselines cover only ten checkpoints, so their coefficients are not matched pairs with the rest. At n≤10 with no interval estimates these are descriptive associations, and single-row differences should not be over-read.
End-to-end
Post-trained
Base checkpoint
pass@1
SWE-Verified
DeepSeek-V4-Pro-Base
0.0
80.6
DeepSeek-V4-Flash-Base
0.0
79.0
Nemotron-Ultra-Base
0.0
70.7
Qwen-3.5-35B-A3B-Base
20.2
69.2
Nemotron-Super-Base
0.0
60.5
Appendix
Table 8: End-to-end evaluation of base checkpoints on SWE-bench Verified, alongside each family’s post-trained score on the same benchmark. All values are pass@1 percentages.
Figure 9: DeepSWE prefix-conditioned pass@K at K=32 has the strongest descriptive agreement with post-trained Terminal-Bench 2.1 performance. Panels use the same ten checkpoints and the same probe settings as fig. 5 ; annotations report Pearson r and Spearman ρ . The BPB axis is reversed so rightward consistently means better.
Source component
Universal representation
Tool definitions
Tool definitions: <JSON>
System message
System: <content>
User message
User: <content>
Assistant text
Assistant: <content>
Assistant tool call
Assistant (tool call): <JSON>
Tool response
Tool Result: <content>
Appendix
Table 9: Universal serialization of trajectory components. Angle-bracketed text denotes the original message content.
DeepSWE
SWE-Verified
Base checkpoint
No think
Think
No think
Think
Hy3-Preview-Base
0.4456
0.5080
0.552
0.583
DeepSeek-V4-Pro-Base
0.4206
0.5026
0.550
0.591
Nemotron-Ultra-Base
0.4235
0.4611
0.538
0.558
DeepSeek-V4-Flash-Base
0.4009
0.4421
0.536
0.553
Kimi-K2-Base
0.4322
0.4047
0.597
0.523
Appendix
Table 10: Source-current four-option Patch MCQ accuracy with cyclic rotation and arithmetic aggregation; chance is 0.250. Both question sets use matched none preambles. Each cell is the one accuracy our source measurements record for that checkpoint and setting; run-to-run spread is not stored with these values ( appendix C ), so column-to-column gaps of a few thousandths should not be read as differences. Bold marks the largest value per column.
DeepSWE
SWE-Verified
Base checkpoint
Argmin
Rotated
Gap
Argmin
Rotated
Gap
Hy3-Preview-Base
0.266
0.508
−0.242
0.359
0.583
−0.224
DeepSeek-V4-Pro-Base
0.283
0.503
−0.220
0.405
0.591
−0.186
Nemotron-Ultra-Base
0.267
0.461
−0.194
0.366
0.558
−0.192
DeepSeek-V4-Flash-Base
0.266
0.442
−0.176
0.380
0.553
−0.173
Kimi-K2-Base
0.279
0.405
−0.126
0.387
0.523
−0.136
Appendix
Table 11: Patch MCQ read ablation on both question sets; chance is 0.250. Argmin-BPB scores each option independently and picks the lowest byte-normalized negative log-likelihood; rotated log-probability is the answer-letter read, presenting all four options jointly under four cyclic rotations, and is reported here at think-1000 to match table 10 . Gap is argmin-BPB minus rotated log-probability, so negative values favor the joint read. Rows are ordered by DeepSWE rotated log-probability.
Base checkpoint
k=1
k=4
Gain
DeepSeek-V4-Pro-Base
0.4555
0.5026
+0.047
Hy3-Preview-Base
0.4271
0.5080
+0.081
Nemotron-Ultra-Base
0.4195
0.4611
+0.042
DeepSeek-V4-Flash-Base
0.3881
0.4421
+0.054
Kimi-K2-Base
0.3741
0.4047
+0.031
Nemotron-Super-Base
0.3727
0.4116
+0.039
Appendix
Table 12: Single-rotation versus full cyclic rotation on the DeepSWE question set, think-1000 arm, over the same ten checkpoints as table 10 ; chance is 0.250. k=1 takes the answer from the first rotation alone and k=4 is the shipped aggregate, both recomputed offline from the same finished runs, so no checkpoint was re-scored; each value is the mean of three rounds. The k=4 column is the DeepSWE think-1000 column of table 10 . Rotation helps on 9 of 10 checkpoints, but not uniformly: the gain spans −0.005 to +0.081 , wider than the gap between adjacent checkpoints in the k=4 ranking, and it reorders the top two. Bold marks the best value per column.
vs. SWE-bench Verified
vs. Terminal-Bench 2.1
Question set
Scoring method
Pearson
Spearman
Pearson
Spearman
DeepSWE
Argmin-BPB
0.601
0.665
0.420
0.382
DeepSWE
Rotated, no think
0.900
0.806
0.686
0.547
DeepSWE
Rotated, think 1000
0.915
0.867
0.871
0.796
SWE-Verified
Argmin-BPB
0.429
0.578
0.359
0.213
SWE-Verified
Rotated, no think
0.909
0.842
0.725
0.578
Appendix
Table 13: Downstream agreement of every Patch MCQ scoring method on the same ten checkpoints, with both Pearson and Spearman correlations. Rows pair a question set with a scoring method; the SWE-Verified question set draws its options from the same benchmark as the first target, so those rows are same-benchmark evidence against it while all rows are cross-domain with respect to Terminal-Bench 2.1. At n=10 with no interval estimates we do not read small differences as separating rows.
Figure 10: Patch MCQ accuracy and downstream agreement, by scoring rule. Panels (a) and (b) rank the same ten checkpoints on each question set; panels (c) and (d) plot the rank agreement tabulated in table 13 against both downstream targets. The joint answer-letter read separates the cohort more sharply than scoring each option independently, and the gap between the two reads is larger than the gap between thinking protocols.
Figure 11: Driving an agent harness from a base checkpoint. (a) The bridge sits between an unmodified agent CLI and the base checkpoint: it renders each agent request into a plain-text transcript, samples a continuation from the checkpoint’s raw completion endpoint, and returns the parsed action to the agent as a native tool call. (b) The prompt at one forward step, assembled as transcript, steer, primer, and prefill. The prefill pre-opens the action envelope so that the checkpoint’s first sampled token already lands inside an open action; only the green span is written by the base checkpoint.
DeepSWE
SWE-bench Verified
Base checkpoint
k=8
k=16
k=32
k=8
k=16
k=32
hy3-base
16.67
20.83
26.67
61.22
70.41
76.53
deepseek-v4-pro-base
37.50
41.67
45.83
68.37
77.55
83.67
ultra-base
12.50
17.50
26.67
58.16
68.37
83.67
dsflash-base
28.33
33.33
40.00
63.27
72.45
79.59
kimi-k2-base
14.17
16.67
21.67
55.10
62.24
68.37
Appendix
Table 14: Prefix-conditioned pass@K on DeepSWE and SWE-bench Verified trajectory prefixes, in percent. DeepSWE uses 120 held-out instances; SWE-bench Verified uses a separately eligibility-filtered cohort of 98 prefixes. Both cover the same ten checkpoints, but the two halves aggregate over different cohort sizes, so their means are not matched pairs. Rollouts fork at the decisive step Ti⋆ : steps before are served verbatim, Ti⋆ is masked, and the checkpoint keeps control for the rest of the episode. An instance counts as resolved if at least one of k continuations passes the verifier. Row order follows table 10 ; bold marks the largest value per column.
vs. SWE-bench Verified
vs. Terminal-Bench 2.1
Probe
n
Pearson
Spearman
Pearson
Spearman
DeepSWE prefix pass@K , k=8
10
0.840
0.915
0.870
0.802
DeepSWE prefix pass@K , k=16
10
0.872
0.952
0.912
0.900
DeepSWE prefix pass@K , k=32
10
0.917
0.951
0.947
0.930
SWE-bench Verified prefix pass@K , k=8
10
0.990
0.988
0.902
0.875
SWE-bench Verified prefix pass@K , k=16
10
0.984
0.988
0.911
0.875
Appendix
Table 15: Downstream validation of prefix-conditioned pass@K against both targets of table 6 , on the same ten-checkpoint cohort used throughout the primary comparisons; higher is better agreement. SWE-bench Verified prefix rows are same-benchmark evidence against the SWE-bench Verified target, not cross-domain validation; both prefix corpora are cross-domain with respect to Terminal-Bench 2.1. The last two rows repeat the canonical all-step/decisive-step comparison from figs. 2 and 7 , correlated against −BPB on the same ten checkpoints. No value is marked best: at n=10 with no bootstrap intervals, we do not read these differences as separating the rows.
Figure 12: Prefix-conditioned pass@K rankings and downstream agreement. Panels (a) and (b) rank the same ten checkpoints under DeepSWE and SWE-bench Verified trajectory prefixes at k∈{8,16,32} , tabulated in table 14 . The two panels are sorted independently and use different x ranges, so marker positions are not comparable across them. Panels (c) and (d) plot the rank agreement of those scores with each downstream target from table 15 ; the dashed line marks Decisive-Action BPB against the same target for reference.
vs. Post-SFT SWE-Verified
vs. Public SWE-Verified
vs. Public SWE-Verified
( n=3 )
( n=3 )
( n=10 )
Evaluation
Pearson
Spearman
Pearson
Spearman
Pearson
Spearman
DeepSWE
Decisive-Action BPB
0.97
0.50
0.96
0.50
0.97
0.96
Patch MCQ
0.91
0.50
0.90
0.50
0.87
0.80
Prefix pass@32
0.97
1.00
0.98
1.00
0.95
0.93
Appendix
Table 16: Comparison of probe correlations across source tasks and post-SFT SWE-bench Verified results for Nemotron 3 Nano, Nemotron 3 Super, and Qwen 3.5 35B A3B. We show correlations against public SWE-bench Verified scores for the same 3 models, as well as all 10 models in our panel.
Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert-built agentic coding tasks closes this gap, and whether what the agent learns transfers beyond the training distribution. We post-train Kimi K2.7 Code, a 1T-parameter (32B active) open-weight mixture-of-experts model, with RL alone on 1,700 tasks: 1,000 repository tasks graded by hidden fail-to-pass tests and by pass-to-pass tests of existing behavior, and 700 terminal tasks graded by expert-written hidden verifiers. The reward is the fraction of target checks passed and drops to zero if any pass-to-pass test fails. One epoch of GSPO on a rank-32 LoRA adapter improves pass@1 on each of the six external benchmarks we evaluated, across three agent harnesses: SWE-Bench Pro (60.1 to 64.8), DeepSWE (31.0 to 43.4), Terminal-Bench 2.1 (67.4 to 82.0), Terminal-Bench 3 (1.4 to 12.1), Terminal-Bench 4 (0.0 to 7.6), and SWE-Marathon (5.0 to 25.0). Pooled over the five independent task sets (Terminal-Bench 4 revises Terminal-Bench 3), the improvement is significant (p < 0.001), and it remains significant on the three sets released after the training data was collected (p = 0.004); the model also improves under both harnesses never used in training. Median trajectories on DeepSWE and Terminal-Bench 3 are 24-35% shorter in agent steps. The base model's failed DeepSWE runs are mostly near-misses, and on the tasks the trained model newly solves, paired trajectories show it avoiding each of the four failure modes above.
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent is structurally isomorphic to a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This conditioning structure exists at internet scale in ordinary code. We exploit it through function-aware fill-in-the-middle (FIM) mid-training: a self-supervised objective that masks functions selected via program dependency graph analysis and a complexity-inferability double criterion. We mid-train Qwen2.5-Coder-Instruct (7B/14B) and Qwen3-8B on a 2.6B-token decontaminated corpus drawn from 968 GitHub repositories, then apply existing agentic post-training pipelines. Mid-training improves SWE-Bench-Verified by +2.8/+3.0 at 7B/14B and by +3.2 on Qwen3-8B; SWE-Bench-Lite gains are +3.7/+4.0/+5.4 on the same models. The improvement holds across two post-training pipelines (R2E-Gym, SWE-Smith) and on a non-Qwen2.5 base (Qwen3-8B with SWE-Lego). Beyond in-domain gains, mid-training also mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding (e.g., LiveCodeBench) and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.
Yubo Wang, Jiarong Liang, Yuxuan Zhang +5
University of Waterloo · 5Vector Institute · University of British Columbia +2
We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests) they can handle before failing. Unlike the prevailing fraction-of-tasks-solved metric, this matches real vibe-coding where sessions run dozens or hundreds of turns. In StaminaBench, agents implement a REST API server and modify it across a tunable number of procedurally generated follow-up change requests - 100 in our experiments, resulting in codebases of up to 6,000 lines. Tests are generated fully programmatically without LLM involvement, ensuring reproducibility and reliability; change sequences are drawn from either a hardcoded or LLM-driven sampler, both constrained to a structured action space to ensure changes are valid. The agent and the server run in an isolated environment and communicate with the benchmark through HTTP, making testing fully black-box and language-agnostic. We evaluate six agent harnesses paired with seven open-source LLMs across 20 scenarios of 100 turns each and find that: (1) all the tested models fail within 5-6 turns, confirming that vibe-coding-style programming without thorough testing produces bugs; (2) passing test feedback back to the agent and allowing it to retry improves passed turn count by up to 12x; and (3) a good harness is required for strong performance: stronger models exhibit up to a 6x gap between their best and worst harness, while weaker models fail with any harness. We release the benchmark and the generated tasks to enable further research into multi-turn coding agent behavior. Benchmark code and data: github.com/amazon-science/StaminaBench.