Better video predictions improve robots’ performance on unfamiliar manipulation tasks

VPP2 combines manipulation-focused video pretraining with action learning while preserving the video model’s ability to generalize.

Top University
Yanjiang Guo · Haodong Yan · Zhide Zhong · Zhongru Zhang · Qingyuan Yang · Qingzhou Lu · +12 more

Robotera · Tsinghua University · HKUST (GZ) · University of California, Berkeley · Shanghai Jiaotong University

Research Digest··2 min read
Guo et al.

The authors continued pretraining a video foundation model on a large, diverse collection of manipulation videos.

Why this paper

From Tsinghua University and 4 others

In one line

Video Prediction Policy 2 achieves strong zero-shot generalization in video prediction and action generation for robot manipulation.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (5 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe