The authors represent each physical process with captions covering relevant entities, causes, interactions, governing principles, temporal progression and effects.
Evolving physics captions improve video models’ physical plausibility
Physis-Lang iteratively refines physical descriptions, retrieves training videos that address model weaknesses, and improves Wan and Cosmos generators across four benchmarks.
Big Tech
Liming Lu · Xianzheng Ma · Wenkun He · Guanqi Zhan · Yilin Zhao · Junyu Chen · +11 more
NVIDIA · University of Oxford · MIT
Research Digest··3 min read
Lu and colleagues test whether explicit natural-language descriptions can carry enough physical knowledge to improve video world models without requiring a separate numerical or latent representation.
Why this paper
From NVIDIA and 2 others
In one line
Using a self-evolving language framework improves video world model physical plausibility without extra visual signals.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§