The authors train causal video diffusion models from scratch without any pretrained bidirectional teacher.
Causal video diffusion matches bidirectional quality without a teacher model
Conditional Residual Prediction reduces history over-reliance, closing a 6-point VBench gap
Chinese Tech
Bowen Zheng · Zhiguang Liu · Jiarong Ou · Rui Chen · Tianyang Hu
The Chinese University of Hong Kong, Shenzhen · Tencent Hunyuan
Research Digest··3 min read
The authors propose Conditional Residual Prediction (CRP), a training method that forces a causal video diffusion model to first predict each frame without using past frames, then let history add only a residual.
Why this paper
From Tencent Hunyuan and The Chinese University of Hong Kong, Shenzhen
In one line
Conditional Residual Prediction reduces causal video models’ dependence on generated history, nearly matching bidirectional quality without a bidirectional teacher.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 2B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§