The authors built WanPE, a 397B-parameter model that expands user requests into plans covering shots, actions, camera movement, lighting and sound.
Video-grounded prompt planning improves long-form text-to-video generation quality
WanPE converts user requests into shot-level cinematic plans while attempting to preserve their meaning across videos lasting up to 30 seconds.
Chinese Tech
Yubo Zhu · Yawen Shao · Ziyun Dai · Zixun Fang · Kai Zhu · Siyang Sun · +24 more
Nanjing University · Wan Team, Alibaba Group · Fudan University · Tsinghua University
Research Digest··2 min read
Zhu et al.
Why this paper
From Wan Team, Alibaba Group and 3 others
In one line
WanPE is a 397B-parameter prompt enhancement model trained on 1.05M real-world videos; it plans cinematic shots and preserves user intent, boosting preference by 10.66-50.86 points.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§