cs.AISep 28, 2026

Improving Large Language Models for Code through Runtime Program-State Reasoning

Authors: Hongwei Li, Spandan Garg, Yufan Huang

Organizations: University of California, Santa Barbara · Microsoft

Abstract

Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

    Sep 23, 2026Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband +3Large Language Model BenchmarksLLM Reasoning Strategies

  2. ProgramBench: Can Language Models Rebuild Programs From Scratch?

    May 5, 2026John Yang, Kilian Lieret, Jeffrey Ma +9Agentic BenchmarksCodebases

  3. StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning

    May 12, 2026Hao Wang, Rui Li, Lei Sha +1Code GenerationSearch-Augmented Reasoning