Visual feedback boosts software engineering agent success rates

New benchmark CUA-SWE shows that agents combining code editing with graphical application interaction significantly outperform code-only approaches.

Top University
Prince Zizhuang Wang · Chenhao Liang · Zelong Xu · Aojie Yuan · Xiaolin Zhou · Haiyue Zhang · +3 more

Carnegie Mellon University · University of Southern California · University of Wisconsin–Madison · Arizona State University · AWS Agentic AI

Research Digest··2 min read
The authors introduce CUA-SWE, a benchmark and evaluation pipeline for software engineering with computer use.

The authors constructed CUA-SWE, a benchmark spanning four software engineering domains (Web, Game, DevOps, Mobile) with tasks that require agents to modify code and configuration, execute commands, interact with running software, and inspect visual feedback.

Why this paper

From Carnegie Mellon University and 4 others · Part of Repository-Level Agent Benchmarks, now 3 papers

In one line

Hybrid agents that combine code editing with GUI interaction outperform code-only agents on software engineering tasks.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.