0 with a dual-stream CrossDiT architecture.
Kandinsky 6.0 jointly generates short videos with synchronized sound
The open models produce five-second clips with speech, lip synchronization and ambient audio from text or image prompts.
Industry
Team Kandinsky · Julia Agafonova · Bulat Akhmatov · Mikhail Aksyutin · Grigorii Alekseenko · Anastasia Aliaskina · +82 more
Kandinsky Lab
Research Digest··2 min read
Team Kandinsky presents two diffusion models that generate video and 44 kHz audio together rather than adding sound in a separate stage.
Why this paper
From Kandinsky Lab
In one line
Kandinsky 6.0 Video generates synchronized 5-second video and 44kHz audio, including lip-sync, from text or image prompts.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§