The authors trained an agent to interact with an environment, letting the student choose actions while a fixed teacher supplied supervision for the states the student actually encountered.
Training agents on their own trajectories accelerates reward-based learning
A teacher-guided warmup improved early reward discovery and subsequent reinforcement learning in interactive agent tasks.
Chinese Tech
Yitong Qiao · Tiantian He · Lei Liu · Yue Shen · Jian Wang · Jinjie Gu · +1 more
Zhejiang University · Ant Healthcare · Ant Group
Research Digest··3 min read
Qiao et al.
Why this paper
From Ant Group and 2 others
In one line
On-policy warmup using teacher supervision on student trajectories accelerates and improves agentic reinforcement learning under sparse rewards.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§