Diagnosing where video models collapse diversity across dimensions

DiVid decomposes video generation diversity into six interpretable axes and reveals that models collapse most severely on temporal factors like motion and camera.

Chinese Tech
Huanran Hu · Zihui Ren · Dingyi Yang · Zhinan Song · Guozheng Wu · Tiezheng Ge · +1 more

Renmin University of China · Alibaba Group · Tianjin University

Research Digest··3 min read
The authors introduce DiVid, a framework that profiles video generation diversity along six interpretable dimensions: Semantic, Style, Subject, Scene, Motion, and Camera.

The authors developed DiVid, a dimension-level diagnostic framework that measures video generation diversity along six complementary axes: Semantic, Style, Subject, Scene, Motion, and Camera.

Why this paper

From Alibaba Group and 2 others

In one line

Video generation models collapse diversity on specific dimensions like Motion and Camera, even when global scores are high.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.