Compression Widens Demographic Fairness Gaps in Whisper Speech Recognition Models

Pruning more than doubles the correction time gap between Black/AA and Asian speakers; quantization causes catastrophic loops on West African accents; distillation narrows gaps in most settings.

Academic
Srishti Ginjala · Eric Fosler-Lussier · Christopher W. Myers · Srinivasan Parthasarathy

The Ohio State University · Air Force Research Laboratory

Research Digest··3 min read
Ginjala et al.

The authors applied three post-training compression methods (Wanda pruning, HQQ and NF4 quantization, and distillation) to the Whisper family (tiny to large-v3) and evaluated demographic fairness across three datasets: FairSpeech, Common Voice 25, and AfriSpeech-200.

Why this paper

From The Ohio State University and Air Force Research Laboratory

In one line

Post-training pruning and quantization can sharply amplify demographic error burdens in Whisper, while distillation usually narrows them.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.