The authors collected YouTube audio using keyword lists for each language, devising methods to improve coverage of medium and low resource languages.
Open YODAS v3 corpus delivers 1.1 million hours of stereo speech
The weakly labeled 48kHz multilingual dataset spans 147 languages and supports high fidelity audio research.
Top University
Carnegie Mellon University · Keio University · Tokyo Metropolitan University · National Institute of Advanced Industrial Science and Technology (AIST)
Research Digest··2 min read
1 million hours of 48kHz stereo audio in 147 languages.
Why this paper
From Carnegie Mellon University and 3 others · Released data
In one line
YODAS v3 is the largest open speech dataset with over 1.1 million hours of 48kHz stereo multilingual audio.
What it released
Data
What we could check
- ·No code link found
- ·No weights link found
- ✓Dataset link in the paper (huggingface.co)
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§