facebook/multilingual_librispeech
Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.
19139k
1---2annotations_creators:3- expert-generated4language_creators:5- crowdsourced6- expert-generated7language:8- de9- nl10- fr11- it12- es13- pt14- pl15- en16license:17- cc-by-4.018multilinguality:19- multilingual20size_categories:21- 100K<n<1M22source_datasets:23- original24task_categories:25- automatic-speech-recognition26- text-to-speech27- text-to-audio28paperswithcode_id: multilingual-librispeech29pretty_name: MultiLingual LibriSpeech30dataset_info:31- config_name: dutch32 features:33 - name: audio34 dtype: audio35 - name: original_path36 dtype: string37 - name: begin_time38 dtype: float6439 - name: end_time40 dtype: float6441 - name: transcript42 dtype: string43 - name: audio_duration44 dtype: float6445 - name: speaker_id46 dtype: string47 - name: chapter_id48 dtype: string49 - name: file50 dtype: string51 - name: id52 dtype: string53 splits:54 - name: dev55 num_bytes: 19995998656 num_examples: 309557 - name: test58 num_bytes: 19929857559 num_examples: 307560 - name: train61 num_bytes: 2393167903162 num_examples: 37428763 - name: 9_hours64 num_bytes: 139884664.66865 num_examples: 215366 - name: 1_hours67 num_bytes: 1546218168 num_examples: 23469 download_size: 2437625662970 dataset_size: 24486284437.66871- config_name: french72 features:73 - name: audio74 dtype: audio75 - name: original_path76 dtype: string77 - name: begin_time78 dtype: float6479 - name: end_time80 dtype: float6481 - name: transcript82 dtype: string83 - name: audio_duration84 dtype: float6485 - name: speaker_id86 dtype: string87 - name: chapter_id88 dtype: string89 - name: file90 dtype: string91 - name: id92 dtype: string93 splits:94 - name: dev95 num_bytes: 157923970.69696 num_examples: 241697 - name: test98 num_bytes: 158352158.58299 num_examples: 2426100 - name: train101 num_bytes: 16984935842.04102 num_examples: 258213103 - name: 9_hours104 num_bytes: 142796680.609105 num_examples: 2167106 - name: 1_hours107 num_bytes: 15675831108 num_examples: 241109 download_size: 17381581776110 dataset_size: 17459684482.927002111- config_name: german112 features:113 - name: audio114 dtype: audio115 - name: original_path116 dtype: string117 - name: begin_time118 dtype: float64119 - name: end_time120 dtype: float64121 - name: transcript122 dtype: string123 - name: audio_duration124 dtype: float64125 - name: speaker_id126 dtype: string127 - name: chapter_id128 dtype: string129 - name: file130 dtype: string131 - name: id132 dtype: string133 splits:134 - name: dev135 num_bytes: 224293581.302136 num_examples: 3469137 - name: test138 num_bytes: 225756069.096139 num_examples: 3394140 - name: train141 num_bytes: 31050881388142 num_examples: 469942143 - name: 9_hours144 num_bytes: 142777983.118145 num_examples: 2194146 - name: 1_hours147 num_bytes: 15714704148 num_examples: 241149 download_size: 31526161821150 dataset_size: 31659423725.516151- config_name: italian152 features:153 - name: audio154 dtype: audio155 - name: original_path156 dtype: string157 - name: begin_time158 dtype: float64159 - name: end_time160 dtype: float64161 - name: transcript162 dtype: string163 - name: audio_duration164 dtype: float64165 - name: speaker_id166 dtype: string167 - name: chapter_id168 dtype: string169 - name: file170 dtype: string171 - name: id172 dtype: string173 splits:174 - name: dev175 num_bytes: 81607596.048176 num_examples: 1248177 - name: test178 num_bytes: 83216752.046179 num_examples: 1262180 - name: train181 num_bytes: 3896742625182 num_examples: 59623183 - name: 9_hours184 num_bytes: 141671904.428185 num_examples: 2173186 - name: 1_hours187 num_bytes: 15560398188 num_examples: 240189 download_size: 4200633596190 dataset_size: 4218799275.522191- config_name: polish192 features:193 - name: audio194 dtype: audio195 - name: original_path196 dtype: string197 - name: begin_time198 dtype: float64199 - name: end_time200 dtype: float64201 - name: transcript202 dtype: string203 - name: audio_duration204 dtype: float64205 - name: speaker_id206 dtype: string207 - name: chapter_id208 dtype: string209 - name: file210 dtype: string211 - name: id212 dtype: string213 splits:214 - name: dev215 num_bytes: 32746725216 num_examples: 512217 - name: test218 num_bytes: 33735044219 num_examples: 520220 - name: train221 num_bytes: 1638889846222 num_examples: 25043223 - name: 9_hours224 num_bytes: 142005461225 num_examples: 2173226 - name: 1_hours227 num_bytes: 15681216228 num_examples: 238229 download_size: 1855342312230 dataset_size: 1863058292231- config_name: portuguese232 features:233 - name: audio234 dtype: audio235 - name: original_path236 dtype: string237 - name: begin_time238 dtype: float64239 - name: end_time240 dtype: float64241 - name: transcript242 dtype: string243 - name: audio_duration244 dtype: float64245 - name: speaker_id246 dtype: string247 - name: chapter_id248 dtype: string249 - name: file250 dtype: string251 - name: id252 dtype: string253 splits:254 - name: dev255 num_bytes: 57533473256 num_examples: 826257 - name: test258 num_bytes: 59141979259 num_examples: 871260 - name: train261 num_bytes: 2518553713.946262 num_examples: 37533263 - name: 9_hours264 num_bytes: 141641902.42265 num_examples: 2116266 - name: 1_hours267 num_bytes: 15697139268 num_examples: 236269 download_size: 2780836500270 dataset_size: 2792568207.366271- config_name: spanish272 features:273 - name: audio274 dtype: audio275 - name: original_path276 dtype: string277 - name: begin_time278 dtype: float64279 - name: end_time280 dtype: float64281 - name: transcript282 dtype: string283 - name: audio_duration284 dtype: float64285 - name: speaker_id286 dtype: string287 - name: chapter_id288 dtype: string289 - name: file290 dtype: string291 - name: id292 dtype: string293 splits:294 - name: dev295 num_bytes: 157804903.144296 num_examples: 2408297 - name: test298 num_bytes: 158526899.32299 num_examples: 2385300 - name: train301 num_bytes: 14562584188302 num_examples: 220701303 - name: 9_hours304 num_bytes: 142473624.48305 num_examples: 2110306 - name: 1_hours307 num_bytes: 15702048308 num_examples: 233309 download_size: 14971394533310 dataset_size: 15037091662.944311configs:312- config_name: dutch313 data_files:314 - split: dev315 path: dutch/dev-*316 - split: test317 path: dutch/test-*318 - split: train319 path: dutch/train-*320 - split: 9_hours321 path: dutch/9_hours-*322 - split: 1_hours323 path: dutch/1_hours-*324- config_name: french325 data_files:326 - split: dev327 path: french/dev-*328 - split: test329 path: french/test-*330 - split: train331 path: french/train-*332 - split: 9_hours333 path: french/9_hours-*334 - split: 1_hours335 path: french/1_hours-*336- config_name: german337 data_files:338 - split: dev339 path: german/dev-*340 - split: test341 path: german/test-*342 - split: train343 path: german/train-*344 - split: 9_hours345 path: german/9_hours-*346 - split: 1_hours347 path: german/1_hours-*348- config_name: italian349 data_files:350 - split: dev351 path: italian/dev-*352 - split: test353 path: italian/test-*354 - split: train355 path: italian/train-*356 - split: 9_hours357 path: italian/9_hours-*358 - split: 1_hours359 path: italian/1_hours-*360- config_name: polish361 data_files:362 - split: dev363 path: polish/dev-*364 - split: test365 path: polish/test-*366 - split: train367 path: polish/train-*368 - split: 9_hours369 path: polish/9_hours-*370 - split: 1_hours371 path: polish/1_hours-*372- config_name: portuguese373 data_files:374 - split: dev375 path: portuguese/dev-*376 - split: test377 path: portuguese/test-*378 - split: train379 path: portuguese/train-*380 - split: 9_hours381 path: portuguese/9_hours-*382 - split: 1_hours383 path: portuguese/1_hours-*384- config_name: spanish385 data_files:386 - split: dev387 path: spanish/dev-*388 - split: test389 path: spanish/test-*390 - split: train391 path: spanish/train-*392 - split: 9_hours393 path: spanish/9_hours-*394 - split: 1_hours395 path: spanish/1_hours-*396---397 398# Dataset Card for MultiLingual LibriSpeech399 400## Table of Contents401- [Dataset Description](#dataset-description)402 - [Dataset Summary](#dataset-summary)403 - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)404 - [Languages](#languages)405 - [How to use](#how-to-use)406- [Dataset Structure](#dataset-structure)407 - [Data Instances](#data-instances)408 - [Data Fields](#data-fields)409 - [Data Splits](#data-splits)410- [Dataset Creation](#dataset-creation)411 - [Curation Rationale](#curation-rationale)412 - [Source Data](#source-data)413 - [Annotations](#annotations)414 - [Personal and Sensitive Information](#personal-and-sensitive-information)415- [Considerations for Using the Data](#considerations-for-using-the-data)416 - [Social Impact of Dataset](#social-impact-of-dataset)417 - [Discussion of Biases](#discussion-of-biases)418 - [Other Known Limitations](#other-known-limitations)419- [Additional Information](#additional-information)420 - [Dataset Curators](#dataset-curators)421 - [Licensing Information](#licensing-information)422 - [Citation Information](#citation-information)423 - [Contributions](#contributions)424 425## Dataset Description426 427- **Homepage:** [MultiLingual LibriSpeech ASR corpus](http://www.openslr.org/94)428- **Repository:** [Needs More Information]429- **Paper:** [MLS: A Large-Scale Multilingual Dataset for Speech Research](https://arxiv.org/abs/2012.03411)430- **Leaderboard:** [🤗 Autoevaluate Leaderboard](https://huggingface.co/spaces/autoevaluate/leaderboards?dataset=facebook%2Fmultilingual_librispeech&only_verified=0&task=automatic-speech-recognition&config=-unspecified-&split=-unspecified-&metric=wer)431 432### Dataset Summary433 434This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. 435The data archives were restructured from the original ones from [OpenSLR](http://www.openslr.org/94) to make it easier to stream.436 437MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 4388 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.439 440### Supported Tasks and Leaderboards441 442- `automatic-speech-recognition`, `speaker-identification`: The dataset can be used to train a model for Automatic Speech Recognition (ASR). The model is presented with an audio file and asked to transcribe the audio file to written text. The most common evaluation metric is the word error rate (WER). The task has an active leaderboard which can be found at https://paperswithcode.com/dataset/multilingual-librispeech and ranks models based on their WER.443- `text-to-speech`, `text-to-audio`: The dataset can also be used to train a model for Text-To-Speech (TTS).444 445### Languages446 447The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish448 449### How to use450 451The `datasets` library allows you to load and pre-process your dataset in pure Python, at scale. The dataset can be downloaded and prepared in one call to your local drive by using the `load_dataset` function. 452 453For example, to download the German config, simply specify the corresponding language config name (i.e., "german" for German):454```python455from datasets import load_dataset456mls = load_dataset("facebook/multilingual_librispeech", "german", split="train")457```458 459Using the datasets library, you can also stream the dataset on-the-fly by adding a `streaming=True` argument to the `load_dataset` function call. Loading a dataset in streaming mode loads individual samples of the dataset at a time, rather than downloading the entire dataset to disk.460```python461from datasets import load_dataset462mls = load_dataset("facebook/multilingual_librispeech", "german", split="train", streaming=True)463print(next(iter(mls)))464```465 466*Bonus*: create a [PyTorch dataloader](https://huggingface.co/docs/datasets/use_with_pytorch) directly with your own datasets (local/streamed).467 468Local:469 470```python471from datasets import load_dataset472from torch.utils.data.sampler import BatchSampler, RandomSampler473mls = load_dataset("facebook/multilingual_librispeech", "german", split="train")474batch_sampler = BatchSampler(RandomSampler(mls), batch_size=32, drop_last=False)475dataloader = DataLoader(mls, batch_sampler=batch_sampler)476```477 478Streaming:479 480```python481from datasets import load_dataset482from torch.utils.data import DataLoader483mls = load_dataset("facebook/multilingual_librispeech", "german", split="train", streaming=True)484dataloader = DataLoader(mls, batch_size=32)485```486 487To find out more about loading and preparing audio datasets, head over to [hf.co/blog/audio-datasets](https://huggingface.co/blog/audio-datasets).488 489### Example scripts490 491Train your own CTC or Seq2Seq Automatic Speech Recognition models on MultiLingual Librispeech with `transformers` - [here](https://github.com/huggingface/transformers/tree/main/examples/pytorch/speech-recognition).492 493## Dataset Structure494 495### Data Instances496 497A typical data point comprises the path to the audio file, usually called `file` and its transcription, called `text`. Some additional information about the speaker and the passage which contains the transcription is provided.498 499```500{'file': '10900_6473_000030.flac',501 'audio': {'path': '10900_6473_000030.flac',502 'array': array([-1.52587891e-04, 6.10351562e-05, 0.00000000e+00, ...,503 4.27246094e-04, 5.49316406e-04, 4.57763672e-04]),504 'sampling_rate': 16000},505 'text': 'więc czego chcecie odemnie spytałem wysłuchawszy tego zadziwiającego opowiadania broń nas stary człowieku broń zakrzyknęli równocześnie obaj posłowie\n',506 'speaker_id': 10900,507 'chapter_id': 6473,508 'id': '10900_6473_000030'}509```510 511 512### Data Fields513 514- file: A filename .flac format.515 516- audio: A dictionary containing the audio filename, the decoded audio array, and the sampling rate. Note that when accessing the audio column: `dataset[0]["audio"]` the audio file is automatically decoded and resampled to `dataset.features["audio"].sampling_rate`. Decoding and resampling of a large number of audio files might take a significant amount of time. Thus it is important to first query the sample index before the `"audio"` column, *i.e.* `dataset[0]["audio"]` should **always** be preferred over `dataset["audio"][0]`.517 518- text: the transcription of the audio file.519 520- id: unique id of the data sample.521 522- speaker_id: unique id of the speaker. The same speaker id can be found for multiple data samples.523- chapter_id: id of the audiobook chapter which includes the transcription.524 525### Data Splits526 527| Number of samples | Train | Train.9h | Train.1h | Dev | Test |528| ----- | ------ | ----- | ---- | ---- | ---- | 529| german | 469942 | 2194 | 241 | 3469 | 3394 |530| dutch | 374287 | 2153 | 234 | 3095 | 3075 |531| french | 258213 | 2167 | 241 | 2416 | 2426 |532| spanish | 220701 | 2110 | 233 | 2408 | 2385 |533| italian | 59623 | 2173 | 240 | 1248 | 1262 |534| portuguese | 37533 | 2116 | 236 | 826 | 871 |535| polish | 25043 | 2173 | 238 | 512 | 520 |536 537## Dataset Creation538 539### Curation Rationale540 541[Needs More Information]542 543### Source Data544 545#### Initial Data Collection and Normalization546 547[Needs More Information]548 549#### Who are the source language producers?550 551[Needs More Information]552 553### Annotations554 555#### Annotation process556 557[Needs More Information]558 559#### Who are the annotators?560 561[Needs More Information]562 563### Personal and Sensitive Information564 565The dataset consists of people who have donated their voice online. You agree to not attempt to determine the identity of speakers in this dataset.566 567## Considerations for Using the Data568 569### Social Impact of Dataset570 571[More Information Needed]572 573### Discussion of Biases574 575[More Information Needed]576 577### Other Known Limitations578 579[Needs More Information]580 581## Additional Information582 583### Dataset Curators584 585[Needs More Information]586 587### Licensing Information588 589Public Domain, Creative Commons Attribution 4.0 International Public License ([CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/legalcode))590 591### Citation Information592 593```594@article{Pratap2020MLSAL,595 title={MLS: A Large-Scale Multilingual Dataset for Speech Research},596 author={Vineel Pratap and Qiantong Xu and Anuroop Sriram and Gabriel Synnaeve and Ronan Collobert},597 journal={ArXiv},598 year={2020},599 volume={abs/2012.03411}600}601```602 603 604### Data Statistics605 606| Duration (h) | Train | Dev | Test |607|--------------|-----------|-------|-------|608| English | 44,659.74 | 15.75 | 15.55 |609| German | 1,966.51 | 14.28 | 14.29 |610| Dutch | 1,554.24 | 12.76 | 12.76 |611| French | 1,076.58 | 10.07 | 10.07 |612| Spanish | 917.68 | 9.99 | 10 |613| Italian | 247.38 | 5.18 | 5.27 |614| Portuguese | 160.96 | 3.64 | 3.74 |615| Polish | 103.65 | 2.08 | 2.14 |616 617| # Speakers | Train | | Dev | | Test | |618|------------|-------|------|-----|----|------|----|619| Gender | M | F | M | F | M | F |620| English | 2742 | 2748 | 21 | 21 | 21 | 21 |621| German | 81 | 95 | 15 | 15 | 15 | 15 |622| Dutch | 9 | 31 | 3 | 3 | 3 | 3 |623| French | 62 | 80 | 9 | 9 | 9 | 9 |624| Spanish | 36 | 50 | 10 | 10 | 10 | 10 |625| Italian | 22 | 43 | 5 | 5 | 5 | 5 |626| Portuguese | 26 | 16 | 5 | 5 | 5 | 5 |627| Polish | 6 | 5 | 2 | 2 | 2 | 2 |628 629| # Hours / Gender | Dev | | Test | |630|------------------|------|------|------|------|631| Gender | M | F | M | F |632| English | 7.76 | 7.99 | 7.62 | 7.93 |633| German | 7.06 | 7.22 | 7 | 7.29 |634| Dutch | 6.44 | 6.32 | 6.72 | 6.04 |635| French | 5.13 | 4.94 | 5.04 | 5.02 |636| Spanish | 4.91 | 5.08 | 4.78 | 5.23 |637| Italian | 2.5 | 2.68 | 2.38 | 2.9 |638| Portuguese | 1.84 | 1.81 | 1.83 | 1.9 |639| Polish | 1.12 | 0.95 | 1.09 | 1.05 |640 641 642 643 644### Contributions645 646Thanks to [@patrickvonplaten](https://github.com/patrickvonplaten) and [@polinaeterna](https://github.com/polinaeterna) for adding this dataset.