The authors developed a data pipeline that generates causally aligned training segments with cross-lingual voice transfer from existing speech translation data.
Causality-aware framework improves simultaneous speech-to-speech translation quality and latency
The authors propose a factorized architecture and adaptive policy that achieve state-of-the-art results with less training data.
Big Tech
Johns Hopkins University · NVIDIA
Research Digest··2 min read
Hussein et al.
Why this paper
From NVIDIA and Johns Hopkins University
In one line
FAST-CAP improves BLEU by up to 1.2 and reduces latency by up to 38.8% in LLM-based simultaneous speech-to-speech translation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§