Synthetic demonstrations let robot policies escape sparse-reward failures

An automated teacher supplied successful manipulation trajectories that enabled subsequent reinforcement learning on tasks where direct PPO received no reward.

Top University
Hiroaki Kingetsu · Hiroaki Kurihara · Kaoru Yokoo · Kenji Fukumizu · Manohar Kaul

Fujitsu Limited · The Institute of Statistical Mathematics

Research Digest··2 min read
The authors introduce SynthDemo-RL, a teacher-student pipeline that generates demonstrations from simulator-privileged state, uses them to fine-tune a vision-language-action policy, and then refines it with reinforcement learning.

Why this paper

From Fujitsu Limited and The Institute of Statistical Mathematics · Part of Agent Self-Improvement, now 10 papers

In one line

SynthDemo-RL generates synthetic demonstrations from simulator state, distills via SFT, and refines with PPO to achieve high success on tasks previously at 0%.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.