NAVA-WAM combines separate video and action diffusion transformers, called Video-DiT and Action-DiT, within a shared attention architecture.
Observation-only videos can directly pretrain robot action policies
NAVA-WAM uses visual-transition prediction to train an action model before adapting it with action-labeled robot demonstrations.
Big Tech
Zhaochong An · Fei Zhang · Menglin Jia · Duncan Frost · Zijian Zhou · Yikai Wang · +7 more
Meta AI · University of Copenhagen · Physical Intelligence · Imperial College London
Research Digest··3 min read
An et al.
Why this paper
From Meta AI and 3 others
In one line
NAVA-WAM directly pretrains an action policy from observation-only videos using future-video flow matching and outperforms prior world action models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§