Depth supervision boosts RGB and geometric consistency in robot world models

A calibration pipeline turns the DROID dataset into metric 3D data, and a Stable Video Diffusion model jointly predicts multi-view RGB and depth, improving RGB by +1.48 dB PSNR.

Academic
Jai Bardhan · Josef Sivic · Vladimir Petrik

Czech Technical University in Prague

Research Digest··2 min read
The authors introduce a calibration pipeline that recalibrates the DROID dataset into a metric 3D corpus (DROID-3D) and train DepthWorld, a video diffusion world model that jointly predicts multi-view RGB and depth.

The authors develop a calibration pipeline that combines learned stereo depth with a joint factor graph optimization, pooling all episodes from the same physical robot to recover shared kinematic parameters and per-scene extrinsics.

Why this paper

From Czech Technical University in Prague

In one line

DepthWorld jointly predicts RGB and depth, improving RGB PSNR by 1.48 dB over an RGB-only baseline while providing metric depth.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe