Training models on runtime program-state reasoning boosts automated software engineering
The authors introduce two complementary program-state reasoning tasks—buggy input-output reasoning and precondition-postcondition reasoning—and incorporate them into a staged post-training pipeline to build Comet-9B, a 9B-parameter language model. They find that adding both tasks to supervised fine-tuning on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified, with further gains from sequential reinforcement learning. Despite its small size, Comet-9B achieves scores comparable to reported results from GPT-5.2 and GPT-4o-based agents.