Open YODAS v3 corpus delivers 1.1 million hours of stereo speech

The weakly labeled 48kHz multilingual dataset spans 147 languages and supports high fidelity audio research.

Top University

Carnegie Mellon University · Keio University · Tokyo Metropolitan University · National Institute of Advanced Industrial Science and Technology (AIST)

Research Digest··2 min read
1 million hours of 48kHz stereo audio in 147 languages.

The authors collected YouTube audio using keyword lists for each language, devising methods to improve coverage of medium and low resource languages.

Why this paper

From Carnegie Mellon University and 3 others · Released data

In one line

YODAS v3 is the largest open speech dataset with over 1.1 million hours of 48kHz stereo multilingual audio.

What it released

Data

What we could check

  • ·No code link found
  • ·No weights link found
  • ✓Dataset link in the paper (huggingface.co)
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.