The authors constructed WideSWE by mining merged pull requests across 103 software ecosystems from the top 200 GitHub organizations.
Coding agents fail at coordinated multi-repository programming tasks
New benchmark WideSWE shows top agent solves only 42.5% of real-world cross-repository tasks, with failures in scope identification and incomplete changes.
Independent
Baoyi Wang · Xingliang Wang · Jinyang Wu · Keming Wu · Chen Zhi · Jianwei Yin
Research Digest··2 min read
The authors introduce WideSWE, a benchmark of 120 real-world software engineering tasks that require coordinated changes across multiple repositories.
Why this paper
Independent
In one line
Coding agents succeed on only 42.50% of cross-repository software tasks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§