New benchmark tests LLMs on repository-level reasoning about code execution

SWE-Flux contains 480 execution-grounded instances across 12 Python repositories, with gold answers automatically harvested from instrumented test runs.

Academic
Hamed Taherkhani · Mohammad Abdollahi · Melika Sepidband · Hridya Dhulipala · Tien N. Nguyen · Hadi Hemmati

York University · The University of Texas at Dallas

Research Digest··2 min read
Taherkhani et al.

The authors constructed a benchmark of 480 questions across 12 real Python repositories.

Why this paper

From York University and The University of Texas at Dallas

In one line

SWE-Flux, a repository-level dynamic execution reasoning benchmark with 480 instances, shows LLMs achieve at most 37% accuracy and struggle with dataflow and inter-procedural state.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.