The authors constructed CUA-SWE, a benchmark spanning four software engineering domains (Web, Game, DevOps, Mobile) with tasks that require agents to modify code and configuration, execute commands, interact with running software, and inspect visual feedback.
Visual feedback boosts software engineering agent success rates
New benchmark CUA-SWE shows that agents combining code editing with graphical application interaction significantly outperform code-only approaches.
Top University
Prince Zizhuang Wang · Chenhao Liang · Zelong Xu · Aojie Yuan · Xiaolin Zhou · Haiyue Zhang · +3 more
Carnegie Mellon University · University of Southern California · University of Wisconsin–Madison · Arizona State University · AWS Agentic AI
Research Digest··2 min read
The authors introduce CUA-SWE, a benchmark and evaluation pipeline for software engineering with computer use.
Why this paper
From Carnegie Mellon University and 4 others · Part of Repository-Level Agent Benchmarks, now 3 papers
In one line
Hybrid agents that combine code editing with GUI interaction outperform code-only agents on software engineering tasks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§