Real-time GUI feedback improves training for computer-use agents

ComputerSD turns observed interface transitions into token-level guidance, while gating that guidance according to whether each action advanced the task.

Chinese Tech
Yong Du · Tongbo Chen · Zhengxi Lu · Yizhou Liu · Bofan Chen · Tao Jiang · +2 more

Zhejiang University · Ant Group

Research Digest··2 min read
Du and colleagues introduce ComputerSD, an online training method that gives computer-use agents feedback after individual GUI actions rather than relying only on final task success.

The authors trained agents in executable computer environments, where actions produce observable GUI transitions.

Why this paper

From Ant Group and Zhejiang University

In one line

ComputerSD improves computer-use agents by turning executed GUI transitions into value-gated token supervision while retaining trajectory-level outcome optimization.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.