Speech & audio

Speech recognition and synthesis, audio understanding, music

9 articles

Speech & audio

Language discrimination training reduces the multilingual gap in speech models

de Seyssel et al. show that strengthening language discrimination during self-supervised pretraining of a bilingual English/French HuBERT model reduces the performance gap compared to monolingual models. On continuous phonetic phone discrimination error, the bilingual baseline of 11.6% dropped to 10.4% (monolingual: 10.8%), while lexical accuracy rose from 52.1% to 56.7% (monolingual: 58.5%) and prosodic lexical performance from 68.9% to 72.9% (monolingual: 72.6%). These gains are largest when the intervention is applied early in training.

3 Oct 2026
Speech & audio

Pruned CTC cuts memory for large-vocabulary ASR training without sacrificing accuracy.

The authors introduce Pruned CTC, a method that reduces the memory footprint of CTC training for large-vocabulary ASR by restricting alignment computation to the subset of vocabulary tokens that appear in each batch, plus the blank token, while retaining full-vocabulary softmax normalization. They prove that this reduction is mathematically equivalent to standard CTC in both loss and gradients. With a Zipformer-M encoder and a 180K vocabulary, Pruned CTC achieves a 5.1x reduction in full-step memory with only a 17% step-time overhead, matching standard CTC accuracy across three corpora.

3 Oct 2026
Speech & audio

Marginal utility allocation of audio tokens improves quality under strict bit budgets

The authors introduce UniAdapt, a method that learns the marginal utility of residual-vector-quantization (RVQ) refinements on a frozen codec and allocates tokens under exact serialized-bit budgets. In utterance-level tests, it reduces Log-STFT distortion by 1.07 to 4.39 percent across speech, music and environmental audio without increasing the budget. A causal streaming variant improves three of four speech rates with zero budget violations across 800 evaluations and runs faster than real time.

3 Oct 2026
Speech & audio

Causality-aware framework improves simultaneous speech-to-speech translation quality and latency

Hussein et al. present FAST-CAP, a causality-aware framework for LLM-based simultaneous speech-to-speech translation. They introduce a factorized architecture (FAST) that separates lexical and acoustic modeling, and a causality-aware adaptive policy (CAP) for read/write decisions. Experiments on Spanish, German, and French show improvements of up to +1.2 BLEU and 26% latency reduction over fixed policies.

3 Oct 2026