Full-duplex speech model scaled to long, multi-party bilingual conversations

The authors extend the Moshi paradigm with a 57.6k-hour synthetic corpus and a new benchmark, training a model that sustains coherent multi-party dialogue over 30+ minutes in English and Chinese.

Top University
Ke Wang · Houxing Ren · Zimu Lu · Yunqiao Yang · Zhuofan Zong · Mingjie Zhan · +1 more

CUHK MMLab · CPII under InnoHK

Research Digest··2 min read
The authors extend the Moshi paradigm to long-horizon, multi-party, bilingual full-duplex speech.

The authors built a data engine that generates parallel-stream audio conversations with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference.

Why this paper

From CUHK MMLab and CPII under InnoHK

In one line

A single end-to-end speech model sustains coherent multi-party bilingual conversations over extended durations and outperforms existing open-source models on MultiTalkBench.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.