The authors observed that in CTC, every valid alignment uses only the target tokens and the blank token; other vocabulary classes contribute only through the softmax normalizer.
Pruned CTC cuts memory for large-vocabulary ASR training without sacrificing accuracy.
The authors prove that restricting CTC alignment to only target tokens plus blank, while preserving full-vocabulary normalization, yields exact loss and gradient equivalence.
Chinese Tech
Yifan Yang · Xiaoyu Yang · Zengrui Jin · Xian Shi · Yuxuan Wang · Yu Xi · +7 more
Shanghai Jiao Tong University · University of Cambridge · Tsinghua University · Alibaba Group · Nankai University
Research Digest··3 min read
The authors introduce Pruned CTC, a method that reduces the memory footprint of CTC training for large-vocabulary ASR by restricting alignment computation to the subset of vocabulary tokens that appear in each batch, plus the blank token, while retaining full-vocabulary softmax normalization.
Why this paper
From Alibaba Group and 5 others
In one line
Pruned CTC restricts CTC alignment computation to target tokens and blank, making large-vocabulary ASR training memory-efficient with exactly equivalent loss and gradients.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 0.6B to 32B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§