Coding agents fail at coordinated multi-repository programming tasks

New benchmark WideSWE shows top agent solves only 42.5% of real-world cross-repository tasks, with failures in scope identification and incomplete changes.

Independent
Baoyi Wang · Xingliang Wang · Jinyang Wu · Keming Wu · Chen Zhi · Jianwei Yin
Research Digest··2 min read
The authors introduce WideSWE, a benchmark of 120 real-world software engineering tasks that require coordinated changes across multiple repositories.

The authors constructed WideSWE by mining merged pull requests across 103 software ecosystems from the top 200 GitHub organizations.

Why this paper

Independent

In one line

Coding agents succeed on only 42.50% of cross-repository software tasks.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.