Counterfactual post-training improves robot world model's action following and physical plausibility

DreamTrue combines offline geometric calibration with reward-guided reinforcement learning on human-annotated defect videos, cutting interaction defect rates from 48.12% to 6.25% on AgiBot.

Chinese Tech
Junyan Li · Ruizhi Li · Yu Liu · Xiangshuo Liu · Mingchao Sun · Hongyu Pan · +3 more

Institute of Automation, Chinese Academy of Sciences · Alibaba Group

Research Digest··2 min read
The authors present DreamTrue, a multi-view robot world model that predicts future video frames conditioned on action sequences.

DreamTrue predicts future observations across camera views and robot embodiments.

Why this paper

From Institute of Automation, Chinese Academy of Sciences and Alibaba Group

In one line

A robot world model using offline geometric calibration and counterfactual post-training improves action following and cuts human-assessed interaction defects from 48.12% to 6.25%.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe