The authors developed InsertionWM, a model-based reinforcement learning approach that learns a visual world model from raw camera observations and proprioceptive data.
A single world model enables robots to insert unseen parts zero-shot
Training on 90 insertion tasks yields 56% zero-shot success, far outperforming model-free baselines.
Big Tech
Nicklas Hansen · Iretiayo Akinola · Yijie Guo · Jie Xu · Bingjie Tang · Hao Su · +4 more
NVIDIA · University of California San Diego · University of Southern California
Research Digest··2 min read
Thread:Robot World Models
Hansen et al.
Why this paper
From NVIDIA and 2 others · Part of Robot World Models, now 3 papers
In one line
A visual world model trained on up to 90 insertion tasks zero-shot assembles unseen objects with 56% success, far above a 7% model-free baseline.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§