On-policy distillation (OPD) trains compact agents by having a teacher score tokens in trajectories sampled from the student's own policy.
Execution graph helps off-the-shelf teachers supervise student agents better
Graph-conditioned on-policy distillation scores student actions against indexed teacher executions, improving success on three agent benchmarks.
Chinese Tech
Xiaohan Yi · Wen Luo · Yani Huang · Junfeng Zhan · Asher Qin · Peilin Zhao · +1 more
Yuanbao Team, Tencent · Tsinghua University · Huazhong University of Science and Technology · Shanghai Jiao Tong University
Research Digest··2 min read
The authors propose GC-OPD, a distillation method that gives an off-the-shelf teacher access to a graph of its own past executions when scoring student trajectories.
Why this paper
From Yuanbao Team, Tencent and 3 others
In one line
Graph-conditioned retrieval of teacher execution histories improves on-policy distillation for compact language agents in multi-turn tasks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§