CoolFace
Datasetpublic

facebook/multilingual_librispeech

Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
191likes39kdownloads
README.md646 linesDownload Raw Back to root
1---2annotations_creators:3- expert-generated4language_creators:5- crowdsourced6- expert-generated7language:8- de9- nl10- fr11- it12- es13- pt14- pl15- en16license:17- cc-by-4.018multilinguality:19- multilingual20size_categories:21- 100K<n<1M22source_datasets:23- original24task_categories:25- automatic-speech-recognition26- text-to-speech27- text-to-audio28paperswithcode_id: multilingual-librispeech29pretty_name: MultiLingual LibriSpeech30dataset_info:31- config_name: dutch32  features:33  - name: audio34    dtype: audio35  - name: original_path36    dtype: string37  - name: begin_time38    dtype: float6439  - name: end_time40    dtype: float6441  - name: transcript42    dtype: string43  - name: audio_duration44    dtype: float6445  - name: speaker_id46    dtype: string47  - name: chapter_id48    dtype: string49  - name: file50    dtype: string51  - name: id52    dtype: string53  splits:54  - name: dev55    num_bytes: 19995998656    num_examples: 309557  - name: test58    num_bytes: 19929857559    num_examples: 307560  - name: train61    num_bytes: 2393167903162    num_examples: 37428763  - name: 9_hours64    num_bytes: 139884664.66865    num_examples: 215366  - name: 1_hours67    num_bytes: 1546218168    num_examples: 23469  download_size: 2437625662970  dataset_size: 24486284437.66871- config_name: french72  features:73  - name: audio74    dtype: audio75  - name: original_path76    dtype: string77  - name: begin_time78    dtype: float6479  - name: end_time80    dtype: float6481  - name: transcript82    dtype: string83  - name: audio_duration84    dtype: float6485  - name: speaker_id86    dtype: string87  - name: chapter_id88    dtype: string89  - name: file90    dtype: string91  - name: id92    dtype: string93  splits:94  - name: dev95    num_bytes: 157923970.69696    num_examples: 241697  - name: test98    num_bytes: 158352158.58299    num_examples: 2426100  - name: train101    num_bytes: 16984935842.04102    num_examples: 258213103  - name: 9_hours104    num_bytes: 142796680.609105    num_examples: 2167106  - name: 1_hours107    num_bytes: 15675831108    num_examples: 241109  download_size: 17381581776110  dataset_size: 17459684482.927002111- config_name: german112  features:113  - name: audio114    dtype: audio115  - name: original_path116    dtype: string117  - name: begin_time118    dtype: float64119  - name: end_time120    dtype: float64121  - name: transcript122    dtype: string123  - name: audio_duration124    dtype: float64125  - name: speaker_id126    dtype: string127  - name: chapter_id128    dtype: string129  - name: file130    dtype: string131  - name: id132    dtype: string133  splits:134  - name: dev135    num_bytes: 224293581.302136    num_examples: 3469137  - name: test138    num_bytes: 225756069.096139    num_examples: 3394140  - name: train141    num_bytes: 31050881388142    num_examples: 469942143  - name: 9_hours144    num_bytes: 142777983.118145    num_examples: 2194146  - name: 1_hours147    num_bytes: 15714704148    num_examples: 241149  download_size: 31526161821150  dataset_size: 31659423725.516151- config_name: italian152  features:153  - name: audio154    dtype: audio155  - name: original_path156    dtype: string157  - name: begin_time158    dtype: float64159  - name: end_time160    dtype: float64161  - name: transcript162    dtype: string163  - name: audio_duration164    dtype: float64165  - name: speaker_id166    dtype: string167  - name: chapter_id168    dtype: string169  - name: file170    dtype: string171  - name: id172    dtype: string173  splits:174  - name: dev175    num_bytes: 81607596.048176    num_examples: 1248177  - name: test178    num_bytes: 83216752.046179    num_examples: 1262180  - name: train181    num_bytes: 3896742625182    num_examples: 59623183  - name: 9_hours184    num_bytes: 141671904.428185    num_examples: 2173186  - name: 1_hours187    num_bytes: 15560398188    num_examples: 240189  download_size: 4200633596190  dataset_size: 4218799275.522191- config_name: polish192  features:193  - name: audio194    dtype: audio195  - name: original_path196    dtype: string197  - name: begin_time198    dtype: float64199  - name: end_time200    dtype: float64201  - name: transcript202    dtype: string203  - name: audio_duration204    dtype: float64205  - name: speaker_id206    dtype: string207  - name: chapter_id208    dtype: string209  - name: file210    dtype: string211  - name: id212    dtype: string213  splits:214  - name: dev215    num_bytes: 32746725216    num_examples: 512217  - name: test218    num_bytes: 33735044219    num_examples: 520220  - name: train221    num_bytes: 1638889846222    num_examples: 25043223  - name: 9_hours224    num_bytes: 142005461225    num_examples: 2173226  - name: 1_hours227    num_bytes: 15681216228    num_examples: 238229  download_size: 1855342312230  dataset_size: 1863058292231- config_name: portuguese232  features:233  - name: audio234    dtype: audio235  - name: original_path236    dtype: string237  - name: begin_time238    dtype: float64239  - name: end_time240    dtype: float64241  - name: transcript242    dtype: string243  - name: audio_duration244    dtype: float64245  - name: speaker_id246    dtype: string247  - name: chapter_id248    dtype: string249  - name: file250    dtype: string251  - name: id252    dtype: string253  splits:254  - name: dev255    num_bytes: 57533473256    num_examples: 826257  - name: test258    num_bytes: 59141979259    num_examples: 871260  - name: train261    num_bytes: 2518553713.946262    num_examples: 37533263  - name: 9_hours264    num_bytes: 141641902.42265    num_examples: 2116266  - name: 1_hours267    num_bytes: 15697139268    num_examples: 236269  download_size: 2780836500270  dataset_size: 2792568207.366271- config_name: spanish272  features:273  - name: audio274    dtype: audio275  - name: original_path276    dtype: string277  - name: begin_time278    dtype: float64279  - name: end_time280    dtype: float64281  - name: transcript282    dtype: string283  - name: audio_duration284    dtype: float64285  - name: speaker_id286    dtype: string287  - name: chapter_id288    dtype: string289  - name: file290    dtype: string291  - name: id292    dtype: string293  splits:294  - name: dev295    num_bytes: 157804903.144296    num_examples: 2408297  - name: test298    num_bytes: 158526899.32299    num_examples: 2385300  - name: train301    num_bytes: 14562584188302    num_examples: 220701303  - name: 9_hours304    num_bytes: 142473624.48305    num_examples: 2110306  - name: 1_hours307    num_bytes: 15702048308    num_examples: 233309  download_size: 14971394533310  dataset_size: 15037091662.944311configs:312- config_name: dutch313  data_files:314  - split: dev315    path: dutch/dev-*316  - split: test317    path: dutch/test-*318  - split: train319    path: dutch/train-*320  - split: 9_hours321    path: dutch/9_hours-*322  - split: 1_hours323    path: dutch/1_hours-*324- config_name: french325  data_files:326  - split: dev327    path: french/dev-*328  - split: test329    path: french/test-*330  - split: train331    path: french/train-*332  - split: 9_hours333    path: french/9_hours-*334  - split: 1_hours335    path: french/1_hours-*336- config_name: german337  data_files:338  - split: dev339    path: german/dev-*340  - split: test341    path: german/test-*342  - split: train343    path: german/train-*344  - split: 9_hours345    path: german/9_hours-*346  - split: 1_hours347    path: german/1_hours-*348- config_name: italian349  data_files:350  - split: dev351    path: italian/dev-*352  - split: test353    path: italian/test-*354  - split: train355    path: italian/train-*356  - split: 9_hours357    path: italian/9_hours-*358  - split: 1_hours359    path: italian/1_hours-*360- config_name: polish361  data_files:362  - split: dev363    path: polish/dev-*364  - split: test365    path: polish/test-*366  - split: train367    path: polish/train-*368  - split: 9_hours369    path: polish/9_hours-*370  - split: 1_hours371    path: polish/1_hours-*372- config_name: portuguese373  data_files:374  - split: dev375    path: portuguese/dev-*376  - split: test377    path: portuguese/test-*378  - split: train379    path: portuguese/train-*380  - split: 9_hours381    path: portuguese/9_hours-*382  - split: 1_hours383    path: portuguese/1_hours-*384- config_name: spanish385  data_files:386  - split: dev387    path: spanish/dev-*388  - split: test389    path: spanish/test-*390  - split: train391    path: spanish/train-*392  - split: 9_hours393    path: spanish/9_hours-*394  - split: 1_hours395    path: spanish/1_hours-*396---397 398# Dataset Card for MultiLingual LibriSpeech399 400## Table of Contents401- [Dataset Description](#dataset-description)402  - [Dataset Summary](#dataset-summary)403  - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)404  - [Languages](#languages)405  - [How to use](#how-to-use)406- [Dataset Structure](#dataset-structure)407  - [Data Instances](#data-instances)408  - [Data Fields](#data-fields)409  - [Data Splits](#data-splits)410- [Dataset Creation](#dataset-creation)411  - [Curation Rationale](#curation-rationale)412  - [Source Data](#source-data)413  - [Annotations](#annotations)414  - [Personal and Sensitive Information](#personal-and-sensitive-information)415- [Considerations for Using the Data](#considerations-for-using-the-data)416  - [Social Impact of Dataset](#social-impact-of-dataset)417  - [Discussion of Biases](#discussion-of-biases)418  - [Other Known Limitations](#other-known-limitations)419- [Additional Information](#additional-information)420  - [Dataset Curators](#dataset-curators)421  - [Licensing Information](#licensing-information)422  - [Citation Information](#citation-information)423  - [Contributions](#contributions)424 425## Dataset Description426 427- **Homepage:** [MultiLingual LibriSpeech ASR corpus](http://www.openslr.org/94)428- **Repository:** [Needs More Information]429- **Paper:** [MLS: A Large-Scale Multilingual Dataset for Speech Research](https://arxiv.org/abs/2012.03411)430- **Leaderboard:** [🤗 Autoevaluate Leaderboard](https://huggingface.co/spaces/autoevaluate/leaderboards?dataset=facebook%2Fmultilingual_librispeech&only_verified=0&task=automatic-speech-recognition&config=-unspecified-&split=-unspecified-&metric=wer)431 432### Dataset Summary433 434This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. 435The data archives were restructured from the original ones from [OpenSLR](http://www.openslr.org/94) to make it easier to stream.436 437MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 4388 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.439 440### Supported Tasks and Leaderboards441 442- `automatic-speech-recognition`, `speaker-identification`: The dataset can be used to train a model for Automatic Speech Recognition (ASR). The model is presented with an audio file and asked to transcribe the audio file to written text. The most common evaluation metric is the word error rate (WER). The task has an active leaderboard which can be found at https://paperswithcode.com/dataset/multilingual-librispeech and ranks models based on their WER.443- `text-to-speech`, `text-to-audio`: The dataset can also be used to train a model for Text-To-Speech (TTS).444 445### Languages446 447The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish448 449### How to use450 451The `datasets` library allows you to load and pre-process your dataset in pure Python, at scale. The dataset can be downloaded and prepared in one call to your local drive by using the `load_dataset` function. 452 453For example, to download the German config, simply specify the corresponding language config name (i.e., "german" for German):454```python455from datasets import load_dataset456mls = load_dataset("facebook/multilingual_librispeech", "german", split="train")457```458 459Using the datasets library, you can also stream the dataset on-the-fly by adding a `streaming=True` argument to the `load_dataset` function call. Loading a dataset in streaming mode loads individual samples of the dataset at a time, rather than downloading the entire dataset to disk.460```python461from datasets import load_dataset462mls = load_dataset("facebook/multilingual_librispeech", "german", split="train", streaming=True)463print(next(iter(mls)))464```465 466*Bonus*: create a [PyTorch dataloader](https://huggingface.co/docs/datasets/use_with_pytorch) directly with your own datasets (local/streamed).467 468Local:469 470```python471from datasets import load_dataset472from torch.utils.data.sampler import BatchSampler, RandomSampler473mls = load_dataset("facebook/multilingual_librispeech", "german", split="train")474batch_sampler = BatchSampler(RandomSampler(mls), batch_size=32, drop_last=False)475dataloader = DataLoader(mls, batch_sampler=batch_sampler)476```477 478Streaming:479 480```python481from datasets import load_dataset482from torch.utils.data import DataLoader483mls = load_dataset("facebook/multilingual_librispeech", "german", split="train", streaming=True)484dataloader = DataLoader(mls, batch_size=32)485```486 487To find out more about loading and preparing audio datasets, head over to [hf.co/blog/audio-datasets](https://huggingface.co/blog/audio-datasets).488 489### Example scripts490 491Train your own CTC or Seq2Seq Automatic Speech Recognition models on MultiLingual Librispeech with `transformers` - [here](https://github.com/huggingface/transformers/tree/main/examples/pytorch/speech-recognition).492 493## Dataset Structure494 495### Data Instances496 497A typical data point comprises the path to the audio file, usually called `file` and its transcription, called `text`. Some additional information about the speaker and the passage which contains the transcription is provided.498 499```500{'file': '10900_6473_000030.flac',501 'audio': {'path': '10900_6473_000030.flac',502  'array': array([-1.52587891e-04,  6.10351562e-05,  0.00000000e+00, ...,503          4.27246094e-04,  5.49316406e-04,  4.57763672e-04]),504  'sampling_rate': 16000},505 'text': 'więc czego chcecie odemnie spytałem wysłuchawszy tego zadziwiającego opowiadania broń nas stary człowieku broń zakrzyknęli równocześnie obaj posłowie\n',506 'speaker_id': 10900,507 'chapter_id': 6473,508 'id': '10900_6473_000030'}509```510 511 512### Data Fields513 514- file: A filename .flac format.515 516- audio: A dictionary containing the audio filename, the decoded audio array, and the sampling rate. Note that when accessing the audio column: `dataset[0]["audio"]` the audio file is automatically decoded and resampled to `dataset.features["audio"].sampling_rate`. Decoding and resampling of a large number of audio files might take a significant amount of time. Thus it is important to first query the sample index before the `"audio"` column, *i.e.* `dataset[0]["audio"]` should **always** be preferred over `dataset["audio"][0]`.517 518- text: the transcription of the audio file.519 520- id: unique id of the data sample.521 522- speaker_id: unique id of the speaker. The same speaker id can be found for multiple data samples.523- chapter_id: id of the audiobook chapter which includes the transcription.524 525### Data Splits526 527|           Number of samples                  | Train | Train.9h | Train.1h  | Dev | Test |528| -----                       | ------ | ----- | ---- | ---- | ---- | 529| german | 469942 | 2194 | 241 | 3469 | 3394 |530| dutch | 374287 | 2153 | 234 | 3095 | 3075 |531| french | 258213 | 2167 | 241 | 2416 | 2426 |532| spanish | 220701 | 2110 | 233 | 2408 | 2385 |533| italian | 59623 | 2173 | 240 | 1248 | 1262 |534| portuguese | 37533 | 2116 | 236 | 826 | 871 |535| polish | 25043 | 2173 | 238 | 512 | 520 |536 537## Dataset Creation538 539### Curation Rationale540 541[Needs More Information]542 543### Source Data544 545#### Initial Data Collection and Normalization546 547[Needs More Information]548 549#### Who are the source language producers?550 551[Needs More Information]552 553### Annotations554 555#### Annotation process556 557[Needs More Information]558 559#### Who are the annotators?560 561[Needs More Information]562 563### Personal and Sensitive Information564 565The dataset consists of people who have donated their voice online. You agree to not attempt to determine the identity of speakers in this dataset.566 567## Considerations for Using the Data568 569### Social Impact of Dataset570 571[More Information Needed]572 573### Discussion of Biases574 575[More Information Needed]576 577### Other Known Limitations578 579[Needs More Information]580 581## Additional Information582 583### Dataset Curators584 585[Needs More Information]586 587### Licensing Information588 589Public Domain, Creative Commons Attribution 4.0 International Public License ([CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/legalcode))590 591### Citation Information592 593```594@article{Pratap2020MLSAL,595  title={MLS: A Large-Scale Multilingual Dataset for Speech Research},596  author={Vineel Pratap and Qiantong Xu and Anuroop Sriram and Gabriel Synnaeve and Ronan Collobert},597  journal={ArXiv},598  year={2020},599  volume={abs/2012.03411}600}601```602 603 604### Data Statistics605 606| Duration (h) | Train     | Dev   | Test  |607|--------------|-----------|-------|-------|608| English      | 44,659.74 | 15.75 | 15.55 |609| German       | 1,966.51  | 14.28 | 14.29 |610| Dutch        | 1,554.24  | 12.76 | 12.76 |611| French       | 1,076.58  | 10.07 | 10.07 |612| Spanish      | 917.68    | 9.99  | 10    |613| Italian      | 247.38    | 5.18  | 5.27  |614| Portuguese   | 160.96    | 3.64  | 3.74  |615| Polish       | 103.65    | 2.08  | 2.14  |616 617| # Speakers | Train |      | Dev |    | Test |    |618|------------|-------|------|-----|----|------|----|619|       Gender   | M     | F    | M   | F  | M    | F  |620| English    | 2742  | 2748 | 21  | 21 | 21   | 21 |621| German     | 81    | 95   | 15  | 15 | 15   | 15 |622| Dutch      | 9     | 31   | 3   | 3  | 3    | 3  |623| French     | 62    | 80   | 9   | 9  | 9    | 9  |624| Spanish    | 36    | 50   | 10  | 10 | 10   | 10 |625| Italian    | 22    | 43   | 5   | 5  | 5    | 5  |626| Portuguese | 26    | 16   | 5   | 5  | 5    | 5  |627| Polish     | 6     | 5    | 2   | 2  | 2    | 2  |628 629| # Hours / Gender | Dev  |      | Test |      |630|------------------|------|------|------|------|631|       Gender   | M    | F    | M    | F    |632| English          | 7.76 | 7.99 | 7.62 | 7.93 |633| German           | 7.06 | 7.22 | 7    | 7.29 |634| Dutch            | 6.44 | 6.32 | 6.72 | 6.04 |635| French           | 5.13 | 4.94 | 5.04 | 5.02 |636| Spanish          | 4.91 | 5.08 | 4.78 | 5.23 |637| Italian          | 2.5  | 2.68 | 2.38 | 2.9  |638| Portuguese       | 1.84 | 1.81 | 1.83 | 1.9  |639| Polish           | 1.12 | 0.95 | 1.09 | 1.05 |640 641 642 643 644### Contributions645 646Thanks to [@patrickvonplaten](https://github.com/patrickvonplaten) and [@polinaeterna](https://github.com/polinaeterna) for adding this dataset.