The authors trained agents in executable computer environments, where actions produce observable GUI transitions.
Real-time GUI feedback improves training for computer-use agents
ComputerSD turns observed interface transitions into token-level guidance, while gating that guidance according to whether each action advanced the task.
Chinese Tech
Yong Du · Tongbo Chen · Zhengxi Lu · Yizhou Liu · Bofan Chen · Tao Jiang · +2 more
Zhejiang University · Ant Group
Research Digest··2 min read
Du and colleagues introduce ComputerSD, an online training method that gives computer-use agents feedback after individual GUI actions rather than relying only on final task success.
Why this paper
From Ant Group and Zhejiang University
In one line
ComputerSD improves computer-use agents by turning executed GUI transitions into value-gated token supervision while retaining trajectory-level outcome optimization.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§