Improving Large Language Models for Code through Runtime Program-State Reasoning
Organizations: University of California, Santa Barbara · Microsoft
Abstract
Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.
Figures & tables
| Model / training stage | Agent | Python | Non-Python | All |
|---|---|---|---|---|
| Public leaderboard baselines | ||||
| DeepSeek V3.2 | SWE-Agent | – | – | 15.56 |
| GPT-OSS-120B | SWE-Agent | 16.20 | ||
| Qwen3-235B-A22B | SWE-Agent | – | – | 21.41 |
| Kimi K2 Instruct | SWE-Agent | 27.67 | ||
| GPT-5.2 | SWE-Agent | – | – | 29.94 |
| Model / training stage | Agent | Success (%) |
| Published leaderboard baselines | ||
| Claude 3.5 Sonnet | OpenHands | 27.7 |
| GPT-4o | AssertFlip | 45.5 |
| Amazon Q | Amazon Q Developer Agent | 51.0 |
| GPT-5-mini | OpenHands | 62.4 |
| Qwen baselines | ||
| SFT variants | RL variants | |||||
| Pure SWE SFT | SWE + Buggy-I/O SFT | Mixed SFT | Buggy-I/O RL | Pre–Post RL | Comet-9B | |
| Success (%) | 7.00 | 5.33 | 6.00 | 6.00 | 11.33 | 10.67 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Mean std. |
|---|---|---|---|---|---|---|
| Qwen3.5-9B Base | 70.0 | 67.8 | 71.1 | 72.2 | 66.7 | 69.6 2.1 |
| Qwen3.5-9B Instruct | 73.3 | 75.6 | 68.9 | 75.6 | 73.3 | 73.3 2.4 |
| Mathematics RL (iter. 74) | 80.0 | 74.4 | 74.4 | 76.7 | 73.3 | 75.8 2.4 |
| Code RL (iter. 45) | 76.7 | 71.1 | 74.4 | 76.7 | 75.6 | 74.9 2.1 |
| Model | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Mean std. |
|---|---|---|---|---|---|---|
| Qwen3.5-9B Base | 37.1 | 31.7 | 37.7 | 32.3 | 32.3 | 34.3 2.6 |
| Qwen3.5-9B Instruct | 50.9 | 57.5 | 52.1 | 51.5 | 54.5 | 53.3 2.4 |
| Mathematics RL (iter. 74) | 44.3 | 48.5 | 55.1 | 51.5 | 51.5 | 50.2 3.6 |
| Code RL (iter. 45) | 56.3 | 51.5 | 54.5 | 53.9 | 52.7 | 53.8 1.6 |
| Benchmark | Tasks | Max. steps | Required output |
|---|---|---|---|
| SWE-bench Pro | 731 | 400 | Source-code patch |
| SWT-Bench Verified | 433 | 250 | Regression-test patch |
| CyberGym | 300 | 400 | Security PoC |
| SWE-bench Pro | CyberGym | |||
|---|---|---|---|---|
| Model / training stage | Python | Non-Python | All | |
| Qwen3.5-9B Base | 6/266 | 3/465 | 9/731 | 11/300 |
| Qwen3.5-9B Instruct | 75/266 | 103/465 | 178/731 | 27/300 |
| Pure SWE SFT | 54/266 | 63/465 | 117/731 | 21/300 |
| SWE + Buggy-I/O SFT | 82/266 | 95/465 | 177/731 | 16/300 |
| Mixed SFT | 80/266 | 90/465 | 170/731 | 18/300 |