The authors built a data engine that generates parallel-stream audio conversations with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference.
Full-duplex speech model scaled to long, multi-party bilingual conversations
The authors extend the Moshi paradigm with a 57.6k-hour synthetic corpus and a new benchmark, training a model that sustains coherent multi-party dialogue over 30+ minutes in English and Chinese.
Top University
Ke Wang · Houxing Ren · Zimu Lu · Yunqiao Yang · Zhuofan Zong · Mingjie Zhan · +1 more
CUHK MMLab · CPII under InnoHK
Research Digest··2 min read
Thread:Latent Speech Reasoning
The authors extend the Moshi paradigm to long-horizon, multi-party, bilingual full-duplex speech.
Why this paper
From CUHK MMLab and CPII under InnoHK
In one line
A single end-to-end speech model sustains coherent multi-party bilingual conversations over extended durations and outperforms existing open-source models on MultiTalkBench.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§