Benchmark tests whether coding agents improve governance without breaking repositories

SWE-Prometheus measures agents’ ability to diagnose governance risks, implement prioritized fixes, preserve behavior and provide execution-backed evidence.

Independent
Jiajun Wu · Leixin Sun · Zihan Tan · Yitao Liu · Shuo Li · Jiaru Qian · +6 more
Research Digest··2 min read
The authors introduce a 60-repository benchmark that replaces a predefined software issue with an open-ended governance objective.

SWE-Prometheus gives an agent a fixed repository snapshot and a broad governance brief rather than a human-identified bug.

Why this paper

Independent · Part of Agent Harness Optimization, now 67 papers

In one line

A repository governance benchmark with execution-backed verification shows AI coding agents can improve tests, docs, deps, and security, but artifact-only gains can mask no real improvement.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.