facebook/multilingual_librispeech
Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.
mls_without_script (#15)
Fix streaming mode (#13)
README updates to include code snippets and correct leaderboard URL (#6)
fix tags
Update dataset script (#3)
Fix `license` metadata (#1)
Update multilingual_librispeech.py
Update dataset links
remove absolute url
get local paths to audio files
Remove task templates
Remove task templates
Update for datasets>=2.1
Update for datasets>=2.1
Update multilingual_librispeech.py
Update multilingual_librispeech.py
Fix streaming filenames
Update data/.gitattributes
Fix newlines
Squashed commit of the following:
add the other languages
add polish
initial commit
