We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses. The strongest setting scores 93.0% on shorter-trace final output, 77.0% on longer-trace final output, and 65.5% and 63.5% on the two state tasks; these scores also hold when missing responses count as wrong. Across 2,397 matched Python comparisons with identical source, changing to the longer-trace input yields 528 correct-to-wrong changes and 147 reversals. Source-clustered analyses preserve this accuracy gap, while adjusted Python models give no evidence of a positive incremental association between cumulative state load and error. Changed inputs and checkpoint tasks alter several factors together, so the gaps do not isolate trace length or an internal state-tracking mechanism. The benchmark exposes errors hidden by short-output scores alone.
Figures & tables
Arm
Program and input
Target
short-trace-final
clean source, shorter-trace input
final output
long-trace-final
same source, longer-trace input
final output
inside-loop-state
instrumented source, long input
state inside loop
post-loop-state
instrumented source, long input
state after loop
Table 1: The canonical four-arm design. The two state arms emit compact JSON with the same ordered fields and top-level types. At least one core field must change between checkpoints.
Study
Lang.
Cases
Sources
Graded
CC median [range]
lcb_python
Python
283
283
7,881
16 [3, 212]
lcb_cpp
C++20
35
35
978
20 [6, 95]
classic
Python
30
30
840
7.5 [5, 13]
codecontests
C++
22
22
612
39.5 [13, 128]
networkx
Python
30
1
840
88 [88, 88]
Total
400
371
11,151
Table 2: Benchmark composition. Graded counts completed predictions included in accuracy, out of 11,200 planned. CC is ΩCC on the clean short-arm source. The NetworkX value repeats because its cases share one source.
Setting
Short final
Long final
Inside state
Post state
gpt-5.6-sol-high
93.0/93.0
77.0/77.0
65.5/65.5
63.5/63.5
gpt-5.6-sol-off
46.2/46.2
32.9/32.5
7.8/7.5
12.6/11.8
glm-5.3-high
87.2/87.2
64.5/64.5
36.2/36.2
33.2/33.2
deepseek-v4-high
91.5/91.5
71.0/71.0
52.8/52.8
39.8/39.8
deepseek-v4-off
19.2/19.2
14.5/14.5
0.2/0.2
0.5/0.5
qwen3.8-27b-high
82.0/82.0
44.2/44.2
16.5/16.5
14.5/14.5
Table 3: Accuracy by arm, in percent. Each cell gives completed-response accuracy C/R followed by all-planned accuracy C/N , which counts nonresponses as wrong; N=400 per cell. Both valid- and invalid-envelope correct answers count. The model label deepseek-v4-pro-0813 is shortened to deepseek-v4 in this table.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Completed
Invalid envelope
Accuracy
All planned
deepseek-v4-pro-0813-high
1600
133
63.8%
63.8%
deepseek-v4-pro-0813-off
1600
154
8.6%
8.6%
glm-5.3-high
1600
396
55.3%
55.3%
gpt-5.6-sol-high
1600
18
74.8%
74.8%
gpt-5.6-sol-off
1553
55
25.2%
24.5%
qwen3.8-27b-high
1598
429
39.4%
39.3%
Appendix
Table 4: Completed-response accuracy, invalid outer envelopes, and the all-planned sensitivity. Each setting plans 1,600 cells. All-planned accuracy counts no responses as wrong.
Model
Predictor (per doubling)
Odds ratio
95% interval
p
Trace + peak
Trace
1.049
[0.988, 1.113]
0.118
Trace + peak
Peak state
1.024
[0.964, 1.087]
0.447
Trace + peak + load
Trace
1.205
[1.025, 1.418]
0.024
Trace + peak + load
Peak state
1.183
[1.009, 1.387]
0.039
Trace + peak + load
State load
0.859
[0.733, 1.006]
0.059
Appendix
Table 6: Exploratory logistic models for prediction error. Both models include study, arm, and setting intercepts; 95% intervals and p -values use source-cluster uncertainty. Odds ratios are per doubling of the named metric.
Cohort
Settings
Pairs
Sources
Decrement (pp)
95% interval
All Python
All seven
2397
314
15.89
[12.12, 20.10]
All Python
Four high
1372
314
22.81
[18.02, 27.92]
Exclude NetworkX
All seven
2187
313
17.33
[13.95, 20.81]
Exclude NetworkX
Four high
1252
313
24.84
[20.93, 28.83]
Appendix
Table 7: Paired short-minus-long decrement with source-cluster percentile intervals. The NetworkX exclusion sensitivity removes the single source reused for its 30 cases.
Figure 2: Error probability against cumulative state load on the classic study, one panel per setting. Squares are binned error rates. Curves and bands show descriptive logistic fits and 95% Wald intervals. These intervals and the panel-header p -values are not cluster robust.
Figure 3: Accuracy against peak state size on lcb_python , with all four arms pooled. Points are predictions, squares are binned accuracy with Wilson intervals, and dashed lines mark overall accuracy. Panel headers report the predeclared rank statistic and nominal permutation p -value.
Python cohorts
C++ cohorts
Metric
Supported
θ range
Supported
θ range
Ω^StateSize
14 / 15
0.51–0.87
0 / 14
0.32–0.64
Ω^StateLoad
14 / 15
0.48–0.88
4 / 14
0.52–0.70
Ω^Trace
12 / 15
0.51–0.89
9 / 14
0.44–0.80
Appendix
Table 8: Pooled support for dynamic metrics by adapter. A supported row has θ>0.5 and unadjusted p<0.05 . The six inestimable Python settings are all-wrong NetworkX settings. Ranges span all estimable series and are not comparable across language groups.
Response contract
Baseline
Invariant
Difference
Exact output only
24 / 36
26 / 36
+2
Visible rationale, then output
24 / 36
24 / 36
0
Appendix
Table 9: Exploratory invariant pilot, outside the 400-case main grid. Each cell contains 12 programs times three fresh sessions. Only the exact output is graded.
Dyson School of Design Engineering Imperial College London London, United Kingdom · Department of Electrical and Electronic Engineering Imperial College London London, United Kingdom