Evidence-driven harness edits improve frozen GUI agents across benchmarks

GUI-HARVEST diagnoses visual execution failures, converts recurring patterns into runtime edits, and validates whether those edits produce the predicted behavior.

Top University
Geyi Yang · Zikun Qu · Xiang Li · Zhiyong Wang · Min Zhang · Shipei Zeng · +1 more

The Chinese University of Hong Kong, Shenzhen · Tianjin University · Harbin Institute of Technology (Shenzhen) · East China Normal University · Shenzhen Research Institute of Big Data

Research Digest··2 min read
Yang et al.

GUI-HARVEST analyzes repeated attempts at computer-use tasks.

Why this paper

From The Chinese University of Hong Kong, Shenzhen and 4 others · Released code

In one line

GUI-HARVEST automatically improves GUI agent harnesses by diagnosing failures from repeated executions and generating reusable edits, boosting performance across models.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.