Causal video diffusion matches bidirectional quality without a teacher model

Conditional Residual Prediction reduces history over-reliance, closing a 6-point VBench gap

Chinese Tech
Bowen Zheng · Zhiguang Liu · Jiarong Ou · Rui Chen · Tianyang Hu

The Chinese University of Hong Kong, Shenzhen · Tencent Hunyuan

Research Digest··3 min read
The authors propose Conditional Residual Prediction (CRP), a training method that forces a causal video diffusion model to first predict each frame without using past frames, then let history add only a residual.

The authors train causal video diffusion models from scratch without any pretrained bidirectional teacher.

Why this paper

From Tencent Hunyuan and The Chinese University of Hong Kong, Shenzhen

In one line

Conditional Residual Prediction reduces causal video models’ dependence on generated history, nearly matching bidirectional quality without a bidirectional teacher.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 2B)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe