Organizations: Carnegie Mellon University, USA · Keio University, Japan · Tokyo Metropolitan University, Japan · National Institute of Advanced Industrial Science and Technology (AIST), Japan
We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio. We first provide the collection methodology for the corpus, where we introduce new techniques for gathering language-balanced speech data. The effectiveness of our approach is shown by the language distribution of the crawled data: 22 languages in YODAS v3 have over 10K hours and 73 languages have over 5K hours of data. We then conduct extensive analyses on the composition of the data, such as the distribution of languages, audio quality, and transcription quality. Finally, we train baseline speech recognition and neural codec models to show the effectiveness of the dataset. Download at https://huggingface.co/datasets/espnet/yodas3.
Figures & tables
Dataset
Languages
Size
Sampling Rate
Stereo
Labeled
License
Common Voice [ 33 ]
137
0.033M hours
48kHz
✗
✓
CC-0
MLS [ 23 ]
8
0.051M hours
16kHz
✗
✓
CC BY 4.0
FLEURS [ 34 ]
102
0.001M hours
16kHz
✗
✓
CC BY 2.5
VoxLingua107 [ 35 ]
107
0.007M hours
16kHz
✗
✓
CC BY 4.0
Librilight [ 18 ]
1
0.060M hours
16kHz
✗
✗
CC BY 4.0
VoxPopuli [ 15 ]
23
0.400M hours
16kHz
✗
✗
CC-0
Table 1 : A comparison of YODAS v3 with a other open large-scale speech datasets. YODAS v3 is the first public dataset to reach a scale of over 1M hours while supporting high-bandwith stereo data.
Figure 1 : Histogram of video upload date. “YODAS2” is the submission deadline of ASRU 2023.
Metadata
Example
Channels
2
Effective channels
2
Sampling rate
48kHz
Max bandwith
20kHz
Views
67
Language
French
Table 2 : Examples of the metadata for each downloaded video
Figure 2 : Data distribution of the top 50 languages in YODAS v3.
Figure 3 : Distribution of videos by total length (log scale), bucketed into groups by minute.
Figure 4 : YODAS v3 data distribution by estimated effective bandwidth; over 92% of recordings exceed 32 kHz.
Figure 5 : Log-scale distribution of data by effective number of channels in each audio file
Threshold
Eng.
Deu.
Fra.
Jap.
Por.
Rus.
Vie.
0.00
15.9
14.4
17.2
27.7
11.1
15.3
20.6
0.10
15.3
14.6
17.8
29.2
11.5
15.4
21.1
0.20
15.3
15.4
18.8
30.9
12.2
16.0
20.8
0.30
15.6
16.3
21.8
34.5
13.4
16.7
23.1
Table 3 : WERs on different languages in Common Voice of ASR models trained on different data filtering thresholds.
LibriTTS
YODAS v3
Training Data
SR
STOI
SPK
STOI
SPK
Baselines
LibriTTS
16kHz
0.95
0.76
0.74
0.69
LibriTTS
24kHz
0.97
0.83
0.80
0.78
AMUSE
16kHz
0.93
0.74
0.81
0.83
AMUSE
44kHz
0.95
0.82
0.90
0.77
Table 4 : Codec evaluation results. AMUSE is a multi-domain mixture for speech, sounds, and music, used by ESPnet-Codec.