Code & software

Code generation, coding agents, software engineering with AI, program synthesis

23 articles

Code & software

Training models on runtime program-state reasoning boosts automated software engineering

The authors introduce two complementary program-state reasoning tasks—buggy input-output reasoning and precondition-postcondition reasoning—and incorporate them into a staged post-training pipeline to build Comet-9B, a 9B-parameter language model. They find that adding both tasks to supervised fine-tuning on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified, with further gains from sequential reinforcement learning. Despite its small size, Comet-9B achieves scores comparable to reported results from GPT-5.2 and GPT-4o-based agents.

3 Oct·3 min
Code & software

New benchmark tests AI's ability to spot invalid code reviews

The authors introduce CRJudgeBench, a benchmark of 1199 instances built from real pull requests and expert-verified perturbations, assessing whether AI agents can distinguish technically valid code review comments from plausible but invalid ones. They also present Sentinel, a repository-grounded agentic judge trained via iterative action-level learning, which achieves 76.60% accuracy on the test set, substantially outperforming base models and showing that general-purpose LLMs struggle with this task.

30 Sept·2 min
Code & software

Code-linked evidence lets AI software claims expire when dependencies change

Tiwari, Vass and Singh present Assay, a Python system that combines repository indexing with an evidence ledger for AI-assisted software delivery. Each claim about code is tied to a Merkle hash of the relevant dependency cone, meaning the module and everything it depends on, so changes invalidate precisely the claims that may no longer hold. Across five public repositories, this approach reduced unnecessary re-verification while a model-free merge gate blocked all nine scripted adversarial behaviors.

30 Sept·3 min