The authors built on a pre-trained speech language model and trained it in two stages.
Chain-of-thought reasoning aligns speech translation duration with source audio
DuraS2ST plans target wording and phonetic length before generating speech, balancing translation quality and temporal consistency.
Big Tech
Yayue Deng · Dingdong Wang · Yuxuan Hu · Jinyu Li · Yanqing Liu · Yuanyuan Wang · +4 more
The Chinese University of Hong Kong · Microsoft Corporation
Research Digest··2 min read
The authors propose DuraS2ST, a speech-to-speech translation framework that uses chain-of-thought reasoning to explicitly plan the target utterance's content and phonetic duration before synthesizing speech.
Why this paper
From Microsoft Corporation and The Chinese University of Hong Kong
In one line
DuraS2ST turns duration control in speech-to-speech translation into an explicit chain-of-thought planning problem solved with reinforcement learning.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§