The authors study world action models, which jointly denoise future video and the actions associated with it.
Simulator-aligned training makes world models more responsive to actions
CtrlWAM pairs perturbed controls with their simulated visual consequences, improving consistency and controllability in driving and robotics.
Big Tech
Chensheng Peng · Wenhao Ding · Ran Tian · Zewei Zhou · Jef Packer · Maximilian Igl · +6 more
NVIDIA · UC Berkeley · UCLA
Research Digest··3 min read
Peng et al.
Why this paper
From NVIDIA and 2 others
In one line
CtrlWAM aligns perturbed actions with their visual consequences in a simulator to improve intent-foresight consistency in world action models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§