datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peoples_speech
Dataset Card for People's Speech
Dataset Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.unsupervised_peoples_speech
Dataset Card for Unsupervised Peoples Speech
Dataset Description
Dataset Summary
The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers.
Point of Contact: MLCommons Datasets Discord
Dataset Structure
This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.peoples_speech_v1.0
Dataset Card for People's Speech
Dataset Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech_v1.0.mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.LHCP-ASR
LHCP-ASR
This dataset is another version of the LHCP-ASR corpus, an English speech dataset for narrow-domain ASR benchmarking in high-energy physics. Unlike the original distribution, which includes video, slides and text data, this version focuses entirely on audio-text pairs
DESCRIPTION
The speech data are 30 hours of LHCP plenary conference talks (2020, 2022) with manual (human) verbatim transcriptions and 205 hours of LHCP conference talks (2020-2022) with automatic… See the full description on the dataset page: https://huggingface.co/datasets/mllp/LHCP-ASR.mls-eng-speaker-descriptions
Dataset Card for Annotations of English MLS
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.indic_tts_ml
Indic TTS Malayalam Speech Corpus
The Malayalam subset of Indic TTS Corpus, taken from
this Kaggle database. The corpus contains
one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given
in the repository.
masri_synthetic
Dataset Card for masri_synthetic
Dataset Summary
The MASRI-SYNTHETIC is a corpus made out of synthesized speech in Maltese. The text-to-speech (TTS) system utilized to produce the utterances was developed by the Research & Development Department of Crimsonwing p.l.c.
The sentences used to create the corpus were extracted from the MLRS Corpus, which is a corpus of written or transcribed Maltese divided into different genres, including: culture, news, academic, religion… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/masri_synthetic.MLC-SLM-Eval
Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) Eval Groundtruth
🖥️ Overview
In the MLC-SLM challenge, we only provided the participants with the audio files of the Eval sets.
Now, we release the oracle segmentation, speaker labels, and transcriptions of the Eval sets to facilitate further research by all participants on the MLC-SLM dataset!
In addition, the MLC-SLM challenge summary paper "Summary on The Multilingual Conversational Speech… See the full description on the dataset page: https://huggingface.co/datasets/bsmu/MLC-SLM-Eval.mls-eng-128kb
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/ntt123/mls-eng-128kb.mls-enhanced-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-10k-tags_tagged_10k_generated.LHCP-ASR-segments
LHCP-ASR Segments
This dataset is a segment-level distribution derived from mllp/LHCP-ASR (and the original LHCP-ASR repository), an English speech corpus for narrow-domain ASR benchmarking in high-energy particle physics.
Unlike previous versions, this repository provides audio directly at the segment level (<30 seconds each) for evaluation, adds talk-level metadata and cleans up transcription tags.
Differences from mllp/LHCP-ASR
No subsets: Directly formatted… See the full description on the dataset page: https://huggingface.co/datasets/mllp/LHCP-ASR-segments.mlx-omni-lora-stt-tts-demoWill be used in the development of the trainer backend of mlx-omni by Neywa Labs.
MLD-VC
🎥 MLD-VC: Multimodal Dataset for Video Conferencing
When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse (CVPR 2026)
📄 [Paper] | 🤗 [Hugging Face Dataset]
📌 Overview
MLD-VC is the first multimodal dataset specifically designed for Audio-Visual Speech Recognition (AVSR) in real-world video conferencing (VC) scenarios.
Unlike traditional AVSR datasets collected in controlled offline environments, MLD-VC… See the full description on the dataset page: https://huggingface.co/datasets/nccm2p2/MLD-VC.Shrutilipi-ML
Shrutilipi-ML Dataset
Overview
This dataset is a specific subset of the Shrutilipi ASR (Automatic Speech Recognition) corpus, containing only the Malayalam language data.
Shrutilipi is a large-scale multilingual speech dataset for Indian languages, originally curated by AI4Bharat. This repository aims to provide a lightweight, language-specific version for researchers and developers focusing on Malayalam speech technology.
Dataset Details
Source… See the full description on the dataset page: https://huggingface.co/datasets/trysem/Shrutilipi-ML.mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/mls-eng-10k-tags_tagged_10k_generated.mls-mimi-codes
Multilingual LibriSpeech (MLS) — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens
for Multilingual LibriSpeech —
LibriVox audiobooks in 7 non-English languages.
English is intentionally excluded. For English Mimi codes, use:
shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits)
shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native)
shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents)
shangeth/jenny-mimi-codes — Jenny… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.whisper-transcripts-ml-street-talk
Dataset Card for "whisper-transcripts-mlst"
More Information needed
ulca_ml
ULCA ASR Dataset Malayalam Speech Corpus
The labelled Malayalam speech subcorpus from the larger ULCA ASR Corpus.
The speech is taken from news broadcasts, and is largely composed of short soundbites with some longer outliers.
ml_superb_br
Description
Partie en breton du jeu de données ML-SUPERB 1.0.En pratique nous nous sommes basés sur espnet/ml_superb_hf.Selon nos estimations, ce jeu de données représente 1 h 10 min et 7s.
Citation
@misc{shi2025mlsuperbmultilingualspeechuniversal,
title={ML-SUPERB: Multilingual Speech Universal PERformance Benchmark},
author={Jiatong Shi and Dan Berrebbi and William Chen and Ho-Lam Chung and En-Pei Hu and Wei Ping Huang and Xuankai Chang and Shang-Wen Li and… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/ml_superb_br.stt-mls-test-fr
MLS — French test split
Split test de Multilingual LibriSpeech (MLS), locale fr, empaqueté en
Parquet shardé avec audio FLAC embarqué — prêt pour load_dataset.
Usage principal : benchmark ASR français (WER / CER) sur livres audio
LibriVox (domaine public).
Contenu
2426 utterances
Audio : FLAC 16 kHz mono (tel que fourni par MLS upstream)
Langue : français (fr)
Licence : CC-BY-4.0 (héritée de MLS / OpenSLR 94)
Durée totale : 10.07 h
Colonnes
Colonne
Type… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-mls-test-fr.mls-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-dacvae.festvox-iiith-mlMLDSUM_NEWmls
Dataset Card for "mls"
More Information needed
espnet_ml_superb_hfCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données espnet/ml_superb_hf.
stt-mls-test
MLS — French test split
Split test de Multilingual LibriSpeech (MLS), locale fr, empaqueté en
Parquet shardé avec audio FLAC embarqué — prêt pour load_dataset.
Usage principal : benchmark ASR français (WER / CER) sur livres audio
LibriVox (domaine public).
Contenu
2426 utterances
Audio : FLAC 16 kHz mono (tel que fourni par MLS upstream)
Langue : français (fr)
Licence : CC-BY-4.0 (héritée de MLS / OpenSLR 94)
Durée totale : 10.07 h
Colonnes
Colonne
Type… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-mls-test.
