SWE-Prometheus gives an agent a fixed repository snapshot and a broad governance brief rather than a human-identified bug.
Benchmark tests whether coding agents improve governance without breaking repositories
SWE-Prometheus measures agents’ ability to diagnose governance risks, implement prioritized fixes, preserve behavior and provide execution-backed evidence.
Independent
Jiajun Wu · Leixin Sun · Zihan Tan · Yitao Liu · Shuo Li · Jiaru Qian · +6 more
Research Digest··2 min read
Thread:Agent Harness Optimization
The authors introduce a 60-repository benchmark that replaces a predefined software issue with an open-ended governance objective.
Why this paper
Independent · Part of Agent Harness Optimization, now 67 papers
In one line
A repository governance benchmark with execution-backed verification shows AI coding agents can improve tests, docs, deps, and security, but artifact-only gains can mask no real improvement.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§