The authors developed VOMMI to connect portable RGB demonstrations to vision-language-action (VLA) post-training.
Portable RGB demonstrations can train mobile manipulation policies and reduce base-velocity errors
VOMMI uses visual odometry refinement and action-group residuals to achieve 18.2% lower base error than robot-collected data while maintaining comparable end-effector accuracy.
Chinese Tech
Yutian Zhang · Xingrui Xiong · Siyuan Ma · Yang Li · Jiawen Wen · Jiaqi Zhai · +9 more
Zhejiang University · Shanghai AI Lab · DeepRobotics · Yale University · Tsinghua University
Research Digest··2 min read
Yutian Zhang et al.
Why this paper
From Shanghai AI Lab and 7 others
In one line
Portable RGB demonstrations via VOMMI train mobile manipulation policies that lower base-velocity error by 18.2% compared to robot-collected demonstrations.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§