The authors developed GUI-SD-v2, a training framework in which a policy learns from a privileged version of itself that receives additional information unavailable during ordinary deployment.
Self-distillation trains GUI agents for longer, memory-dependent tasks
GUI-SD-v2 uses privileged reasoning and memory guidance during training to improve multi-step interaction with mobile interfaces.
Chinese Tech
Yan Zhang · Daiqing Wu · Huawen Shen · Liang Li · Gang Cao · Zhi Gong · +4 more
Institute of Information Engineering, Chinese Academy of Sciences · Tencent · Tsinghua University · University of Chinese Academy of Sciences · Nankai University
Research Digest··2 min read
Thread:Agent Self-Improvement
Zhang et al.
Why this paper
From Institute of Information Engineering, Chinese Academy of Sciences and 4 others · Part of Agent Self-Improvement, now 18 papers
In one line
GUI-SD-v2 extends on-policy self-distillation to multi-turn GUI agents, outperforming state-of-the-art methods.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§