Organizations: Carnegie Mellon University, USA · Keio University, Japan · Tokyo Metropolitan University, Japan · National Institute of Advanced Industrial Science and Technology (AIST), Japan
We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio. We first provide the collection methodology for the corpus, where we introduce new techniques for gathering language-balanced speech data. The effectiveness of our approach is shown by the language distribution of the crawled data: 22 languages in YODAS v3 have over 10K hours and 73 languages have over 5K hours of data. We then conduct extensive analyses on the composition of the data, such as the distribution of languages, audio quality, and transcription quality. Finally, we train baseline speech recognition and neural codec models to show the effectiveness of the dataset. Download at https://huggingface.co/datasets/espnet/yodas3.
Figures & tables
Dataset
Languages
Size
Sampling Rate
Stereo
Labeled
License
Common Voice [ 33 ]
137
0.033M hours
48kHz
✗
✓
CC-0
MLS [ 23 ]
8
0.051M hours
16kHz
✗
✓
CC BY 4.0
FLEURS [ 34 ]
102
0.001M hours
16kHz
✗
✓
CC BY 2.5
VoxLingua107 [ 35 ]
107
0.007M hours
16kHz
✗
✓
CC BY 4.0
Librilight [ 18 ]
1
0.060M hours
16kHz
✗
✗
CC BY 4.0
VoxPopuli [ 15 ]
23
0.400M hours
16kHz
✗
✗
CC-0
Table 1 : A comparison of YODAS v3 with a other open large-scale speech datasets. YODAS v3 is the first public dataset to reach a scale of over 1M hours while supporting high-bandwith stereo data.
Figure 1 : Histogram of video upload date. “YODAS2” is the submission deadline of ASRU 2023.
Metadata
Example
Channels
2
Effective channels
2
Sampling rate
48kHz
Max bandwith
20kHz
Views
67
Language
French
Table 2 : Examples of the metadata for each downloaded video
Figure 2 : Data distribution of the top 50 languages in YODAS v3.
Figure 3 : Distribution of videos by total length (log scale), bucketed into groups by minute.
Figure 4 : YODAS v3 data distribution by estimated effective bandwidth; over 92% of recordings exceed 32 kHz.
Figure 5 : Log-scale distribution of data by effective number of channels in each audio file
Threshold
Eng.
Deu.
Fra.
Jap.
Por.
Rus.
Vie.
0.00
15.9
14.4
17.2
27.7
11.1
15.3
20.6
0.10
15.3
14.6
17.8
29.2
11.5
15.4
21.1
0.20
15.3
15.4
18.8
30.9
12.2
16.0
20.8
0.30
15.6
16.3
21.8
34.5
13.4
16.7
23.1
Table 3 : WERs on different languages in Common Voice of ASR models trained on different data filtering thresholds.
LibriTTS
YODAS v3
Training Data
SR
STOI
SPK
STOI
SPK
Baselines
LibriTTS
16kHz
0.95
0.76
0.74
0.69
LibriTTS
24kHz
0.97
0.83
0.80
0.78
AMUSE
16kHz
0.93
0.74
0.81
0.83
AMUSE
44kHz
0.95
0.82
0.90
0.77
Table 4 : Codec evaluation results. AMUSE is a multi-domain mixture for speech, sounds, and music, used by ESPnet-Codec.
Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sharply for most languages due to limited publicly available aligned data. To this end, we introduce WorldSpeech, a 24 kHz multilingual speech corpus comprising 65k hours of aligned audio-transcript data across 76 languages, collected from diverse public sources including parliamentary proceedings, international broadcasts, and public-domain audiobooks. For 37 languages, WorldSpeech provides more than 200 hours of aligned speech, with 28 exceeding 500 hours and 24 surpassing 1k hours. Fine-tuning existing ASR models on WorldSpeech results in an average relative Word-Error-Rate reduction of 63.5% across 11 typologically diverse languages.
Antonis Asonitis, Luca A. Lanzendörfer, Frédéric Berdoz +1
We present CS-YODAS, a Creative Commons-licensed dataset of in-the-wild code-switched speech mined from multilingual YouTube data. Code-switching (CS), or the alternation between languages within an utterance or conversation, is common in multilingual settings but remains underrepresented in existing CS speech resources, which are typically small, domain-specific, or artificially constructed. Building on the YODAS corpus, we develop a scalable, human-in-the-loop pipeline for identifying and validating naturally occurring code-switching. The resulting dataset, which totals 313 hours and spans 7 matrix languages, provides diverse, real-world examples of spontaneous code-switched speech. We further analyze the distribution and characteristics of code-switching in the wild, examining language-pair frequencies and switching patterns, and report baseline results for spoken language identification. We hope that CS-YODAS will encourage broader and more comprehensive research on code-switched speech. Dataset link: https://huggingface.co/datasets/byan/cs-yodas.
Brian Yan, Qingzheng Wang, Matthew Wiesner +9
Carnegie Mellon University · Johns Hopkins University · University of Texas at Austin +4
Automatic speech recognition (ASR) has advanced remarkably for standard speech, yet speech affected by neurological conditions remains a challenge. We present S-DiverSe (Spanish Diverse Speech), a corpus of 3.2 hours of in-the-wild Spanish speech from 22 speakers with amyotrophic lateral sclerosis, Parkinson's disease, and stroke. The dataset contains 444 manually transcribed audio segments with metadata on speaker sex, disease type, and intelligibility. S-DiverSe is designed to support ASR evaluation and development for neurologically affected Spanish speech. We describe the dataset, analyze its composition, and report baseline ASR results alongside initial adaptation experiments. Our findings reveal that heuristic text post-processing is more robust than fine-tuning for out-of-domain neurological Spanish speech. This underscores the need for dedicated in-the-wild Spanish benchmarks.
Fernando López, Fernando Ibañez, Ana Martínez +4
Scientific Research, Telefónica Innovación Digital, Spain · Universidad Autónoma de Madrid, Spain · Brno University of Technology, Czech Republic