Models struggle to predict software behavior after stateful interventions

CTE-Bench tests whether models can forecast 40 service responses after code or stored-state changes, revealing heavy dependence on correct intermediate feedback.

Independent
Xinran Zhang

Independent Researcher

Research Digest··2 min read
Xinran Zhang introduces CTE-Bench, an executable benchmark that isolates whether models can predict the downstream effects of modifying a stateful software service.

CTE-Bench-Core-v1 contains 255 scenarios spanning six deterministic Python services.

Why this paper

From Independent Researcher · Released data

In one line

Current models predict stateful service changes poorly without correct intermediate feedback, with rollout errors compounding across later calls.

What it released

Data

What we could check

  • ·No code link found
  • ·No weights link found
  • ✓Dataset link in the paper (huggingface.co)
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.