The authors introduce a method that converts visual observations into an embodiment-agnostic representation by first inferring a subtask and an interaction triplet (gripper, held object, target object) using a vision-language model (VLM).
Canonical gripper-frame representation achieves zero-shot cross-embodiment transfer for two-finger manipulation
By decoupling task semantics from hardware-specific visual features, the framework allows policies trained on one robot to generalize to completely different platforms without retraining.
Chinese Tech
Renmin University of China · Zhipu AI · Tsinghua University
Research Digest··2 min read
Thread:Robot World Models
The authors propose an interaction-centric framework that uses a parameterized universal gripper abstraction to transform RGB-D observations into a canonical gripper frame.
Why this paper
From Zhipu AI and 2 others · Part of Robot World Models, now 3 papers
In one line
Interaction-centric representation with a canonical gripper frame enables cross-embodiment zero-shot transfer for two-finger gripper manipulation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§