CTE-Bench-Core-v1 contains 255 scenarios spanning six deterministic Python services.
Models struggle to predict software behavior after stateful interventions
CTE-Bench tests whether models can forecast 40 service responses after code or stored-state changes, revealing heavy dependence on correct intermediate feedback.
Independent
Xinran Zhang
Independent Researcher
Research Digest··2 min read
Xinran Zhang introduces CTE-Bench, an executable benchmark that isolates whether models can predict the downstream effects of modifying a stateful software service.
Why this paper
From Independent Researcher · Released data
In one line
Current models predict stateful service changes poorly without correct intermediate feedback, with rollout errors compounding across later calls.
What it released
Data
What we could check
- ·No code link found
- ·No weights link found
- ✓Dataset link in the paper (huggingface.co)
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§