12 papers this week in Speech & audio12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Speech & audio research

Latest Paper· Speech & audio

VoiceNet tests recognition of nuanced emotions and speaking styles

Schuhmann and colleagues introduce VoiceNet, a benchmark for recognizing fine-grained emotion and performance characteristics in permissively licensed, real-world speech. Their VoiceCLAP models substantially outperform existing audio-text baselines, although the evaluation measures representations and retrieval rather than conversational understanding.

Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby +7
Speech & audio

Language discrimination training reduces the multilingual gap in speech models

de Seyssel et al. show that strengthening language discrimination during self-supervised pretraining of a bilingual English/French HuBERT model reduces the performance gap compared to monolingual models. On continuous phonetic phone discrimination error, the bilingual baseline of 11.6% dropped to 10.4% (monolingual: 10.8%), while lexical accuracy rose from 52.1% to 56.7% (monolingual: 58.5%) and prosodic lexical performance from 68.9% to 72.9% (monolingual: 72.6%). These gains are largest when the intervention is applied early in training.

today
Speech & audio

Pruned CTC cuts memory for large-vocabulary ASR training without sacrificing accuracy.

The authors introduce Pruned CTC, a method that reduces the memory footprint of CTC training for large-vocabulary ASR by restricting alignment computation to the subset of vocabulary tokens that appear in each batch, plus the blank token, while retaining full-vocabulary softmax normalization. They prove that this reduction is mathematically equivalent to standard CTC in both loss and gradients. With a Zipformer-M encoder and a 180K vocabulary, Pruned CTC achieves a 5.1x reduction in full-step memory with only a 17% step-time overhead, matching standard CTC accuracy across three corpora.

today
Speech & audio

Marginal utility allocation of audio tokens improves quality under strict bit budgets

The authors introduce UniAdapt, a method that learns the marginal utility of residual-vector-quantization (RVQ) refinements on a frozen codec and allocates tokens under exact serialized-bit budgets. In utterance-level tests, it reduces Log-STFT distortion by 1.07 to 4.39 percent across speech, music and environmental audio without increasing the budget. A causal streaming variant improves three of four speech rates with zero budget violations across 800 evaluations and runs faster than real time.

today
Safety & security

Compression Widens Demographic Fairness Gaps in Whisper Speech Recognition Models

Ginjala et al. examine how post-training compression (pruning, quantization, distillation) of Whisper ASR models redistributes error burden across demographic groups. On FairSpeech, pruning Whisper-large-v3 widens the temporal-taxation differential between Black/AA and Asian speakers by 111%, while INT4 quantization on West African accents increases catastrophic transcript loops by factors of five to seven. The authors conclude that fairness audits on full-precision models do not capture the deployment-time burden compression imposes on already-marginalized speakers.

today

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses