The authors start from a frozen bidirectional audio-video diffusion transformer.
Parallel adapter composition yields real-time joint audio-video generation without chained training
By adding a causal adapter and a few-step adapter to a frozen backbone, the authors achieve real-time streaming at 26 fps with stable quality.
Industry
Jingyu Li · Xiaoxiao Xiang · Yiwen Guo
LIGHTSPEED · Independent Researcher
Research Digest··2 min read
Li et al.
Why this paper
From LIGHTSPEED and Independent Researcher
In one line
Parallel adapter composition achieves real-time streaming audio-video generation by adding separately trained causal and few-step adapters.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (gpus 6)
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§