We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses. The strongest setting scores 93.0% on shorter-trace final output, 77.0% on longer-trace final output, and 65.5% and 63.5% on the two state tasks; these scores also hold when missing responses count as wrong. Across 2,397 matched Python comparisons with identical source, changing to the longer-trace input yields 528 correct-to-wrong changes and 147 reversals. Source-clustered analyses preserve this accuracy gap, while adjusted Python models give no evidence of a positive incremental association between cumulative state load and error. Changed inputs and checkpoint tasks alter several factors together, so the gaps do not isolate trace length or an internal state-tracking mechanism. The benchmark exposes errors hidden by short-output scores alone.
Figures & tables
Arm
Program and input
Target
short-trace-final
clean source, shorter-trace input
final output
long-trace-final
same source, longer-trace input
final output
inside-loop-state
instrumented source, long input
state inside loop
post-loop-state
instrumented source, long input
state after loop
Table 1: The canonical four-arm design. The two state arms emit compact JSON with the same ordered fields and top-level types. At least one core field must change between checkpoints.
Study
Lang.
Cases
Sources
Graded
CC median [range]
lcb_python
Python
283
283
7,881
16 [3, 212]
lcb_cpp
C++20
35
35
978
20 [6, 95]
classic
Python
30
30
840
7.5 [5, 13]
codecontests
C++
22
22
612
39.5 [13, 128]
networkx
Python
30
1
840
88 [88, 88]
Total
400
371
11,151
Table 2: Benchmark composition. Graded counts completed predictions included in accuracy, out of 11,200 planned. CC is ΩCC on the clean short-arm source. The NetworkX value repeats because its cases share one source.
Setting
Short final
Long final
Inside state
Post state
gpt-5.6-sol-high
93.0/93.0
77.0/77.0
65.5/65.5
63.5/63.5
gpt-5.6-sol-off
46.2/46.2
32.9/32.5
7.8/7.5
12.6/11.8
glm-5.3-high
87.2/87.2
64.5/64.5
36.2/36.2
33.2/33.2
deepseek-v4-high
91.5/91.5
71.0/71.0
52.8/52.8
39.8/39.8
deepseek-v4-off
19.2/19.2
14.5/14.5
0.2/0.2
0.5/0.5
qwen3.8-27b-high
82.0/82.0
44.2/44.2
16.5/16.5
14.5/14.5
Table 3: Accuracy by arm, in percent. Each cell gives completed-response accuracy C/R followed by all-planned accuracy C/N , which counts nonresponses as wrong; N=400 per cell. Both valid- and invalid-envelope correct answers count. The model label deepseek-v4-pro-0813 is shortened to deepseek-v4 in this table.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Completed
Invalid envelope
Accuracy
All planned
deepseek-v4-pro-0813-high
1600
133
63.8%
63.8%
deepseek-v4-pro-0813-off
1600
154
8.6%
8.6%
glm-5.3-high
1600
396
55.3%
55.3%
gpt-5.6-sol-high
1600
18
74.8%
74.8%
gpt-5.6-sol-off
1553
55
25.2%
24.5%
qwen3.8-27b-high
1598
429
39.4%
39.3%
Appendix
Table 4: Completed-response accuracy, invalid outer envelopes, and the all-planned sensitivity. Each setting plans 1,600 cells. All-planned accuracy counts no responses as wrong.
Model
Predictor (per doubling)
Odds ratio
95% interval
p
Trace + peak
Trace
1.049
[0.988, 1.113]
0.118
Trace + peak
Peak state
1.024
[0.964, 1.087]
0.447
Trace + peak + load
Trace
1.205
[1.025, 1.418]
0.024
Trace + peak + load
Peak state
1.183
[1.009, 1.387]
0.039
Trace + peak + load
State load
0.859
[0.733, 1.006]
0.059
Appendix
Table 6: Exploratory logistic models for prediction error. Both models include study, arm, and setting intercepts; 95% intervals and p -values use source-cluster uncertainty. Odds ratios are per doubling of the named metric.
Cohort
Settings
Pairs
Sources
Decrement (pp)
95% interval
All Python
All seven
2397
314
15.89
[12.12, 20.10]
All Python
Four high
1372
314
22.81
[18.02, 27.92]
Exclude NetworkX
All seven
2187
313
17.33
[13.95, 20.81]
Exclude NetworkX
Four high
1252
313
24.84
[20.93, 28.83]
Appendix
Table 7: Paired short-minus-long decrement with source-cluster percentile intervals. The NetworkX exclusion sensitivity removes the single source reused for its 30 cases.
Figure 2: Error probability against cumulative state load on the classic study, one panel per setting. Squares are binned error rates. Curves and bands show descriptive logistic fits and 95% Wald intervals. These intervals and the panel-header p -values are not cluster robust.
Figure 3: Accuracy against peak state size on lcb_python , with all four arms pooled. Points are predictions, squares are binned accuracy with Wilson intervals, and dashed lines mark overall accuracy. Panel headers report the predeclared rank statistic and nominal permutation p -value.
Python cohorts
C++ cohorts
Metric
Supported
θ range
Supported
θ range
Ω^StateSize
14 / 15
0.51–0.87
0 / 14
0.32–0.64
Ω^StateLoad
14 / 15
0.48–0.88
4 / 14
0.52–0.70
Ω^Trace
12 / 15
0.51–0.89
9 / 14
0.44–0.80
Appendix
Table 8: Pooled support for dynamic metrics by adapter. A supported row has θ>0.5 and unadjusted p<0.05 . The six inestimable Python settings are all-wrong NetworkX settings. Ranges span all estimable series and are not comparable across language groups.
Response contract
Baseline
Invariant
Difference
Exact output only
24 / 36
26 / 36
+2
Visible rationale, then output
24 / 36
24 / 36
0
Appendix
Table 9: Exploratory invariant pilot, outside the 400-case main grid. Each cell contains 12 programs times three fresh sessions. Only the exact output is graded.
Recent advances in large language models (LLMs) have shown that test-time scaling can substantially improve model performance on complex tasks, particularly in the coding domain. Under this paradigm, models use a larger token budget during inference to generate intermediate reasoning traces before producing a final answer. However, current evaluations primarily rely on competitive programming benchmarks, which may not capture the full range of reasoning abilities. In this work, we perform a systematic study of frontier reasoning models to understand their performance on real-world coding benchmarks. To gain more insights into the performance of such models, we devise a programmatic way to {\em automatically generate} coding tasks of arbitrary difficulty and structure from existing benchmarks. Using this framework, our analysis reveals that the structure of a reasoning trace, not just its contents, is a strong predictor of correctness. Motivated by this, we propose structured thought-trees as means to represent reasoning traces. To illustrate their use, we train a lightweight classifier on features extracted from thought-trees to predict trace correctness, and demonstrate that flagging and retrying structurally anomalous traces based on the extracted features yields consistent gains at lower complexity levels.
Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.
Hongwei Li, Spandan Garg, Yufan Huang
University of California, Santa Barbara · Microsoft
Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation framework that separates scorer-independent execution evidence, including termination, answer exposure, parseability, and completion length, from scorer-dependent correctness. Across 2,550 outputs from five fixed Qwen and DeepSeek configurations on MATH and ARC-Challenge, matched 2,048-token limits produce sharply different execution mixtures: 49 of 450 Qwen MATH outputs terminate without a final answer, compared with 5 of 300 DeepSeek MATH outputs and none of the 750 ARC outputs. Among the same 300 DeepSeek MATH question-model pairs, no missing-final length termination is observed at 8,192 tokens. A coverage-audited targeted verification study further shows that candidate-selection and aggregation policies can substantially alter comparative accuracy estimates. These results demonstrate that accuracy conflates execution case mix with verification policy. Evaluations of test-time methods should therefore report pre-intervention execution states, verification coverage, and scorer provenance alongside accuracy.
Zongyou Yang, Yinghan Hou
Dyson School of Design Engineering Imperial College London London, United Kingdom · Department of Electrical and Electronic Engineering Imperial College London London, United Kingdom