Synthetic demonstrations let robot policies escape sparse-reward failures
An automated teacher supplied successful manipulation trajectories that enabled subsequent reinforcement learning on tasks where direct PPO received no reward.
Top University
Hiroaki Kingetsu · Hiroaki Kurihara · Kaoru Yokoo · Kenji Fukumizu · Manohar Kaul
Fujitsu Limited · The Institute of Statistical Mathematics
Research Digest··2 min read
Thread:Agent Self-Improvement
The authors introduce SynthDemo-RL, a teacher-student pipeline that generates demonstrations from simulator-privileged state, uses them to fine-tune a vision-language-action policy, and then refines it with reinforcement learning.
Why this paper
From Fujitsu Limited and The Institute of Statistical Mathematics · Part of Agent Self-Improvement, now 10 papers
In one line
SynthDemo-RL generates synthetic demonstrations from simulator state, distills via SFT, and refines with PPO to achieve high success on tasks previously at 0%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§