Ding et al.
New benchmark measures LLMs' ability to build complete codebases from scratch
E2E-SWE includes 186 tasks across 11 languages, with pass@1 scores ranging from 11.7% to 67.7% across 13 frontier models.
Industry
Hantian Ding · Chloe Bi · Jiacheng Zhu · John Yang · Matt Deitke · Pengcheng Yin · +2 more
Meta Superintelligence Labs
Research Digest··2 min read
The authors introduce E2E-SWE, a benchmark of 186 whole-repository generation tasks across 11 programming languages.
Why this paper
From Meta Superintelligence Labs
In one line
Whole-repository generation from specs is feasible, with LLM pass@1 ranging from 11.7% to 67.7% on the E2E-SWE benchmark.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§