The authors constructed a benchmark of 480 questions across 12 real Python repositories.
New benchmark tests LLMs on repository-level reasoning about code execution
SWE-Flux contains 480 execution-grounded instances across 12 Python repositories, with gold answers automatically harvested from instrumented test runs.
Academic
Hamed Taherkhani · Mohammad Abdollahi · Melika Sepidband · Hridya Dhulipala · Tien N. Nguyen · Hadi Hemmati
York University · The University of Texas at Dallas
Research Digest··2 min read
Taherkhani et al.
Why this paper
From York University and The University of Texas at Dallas
In one line
SWE-Flux, a repository-level dynamic execution reasoning benchmark with 480 instances, shows LLMs achieve at most 37% accuracy and struggle with dataflow and inter-procedural state.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§