datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.ghana-english-asr-2700hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.ghana-english-speech-600hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-600hrs.translated-german-english-asr
Translated German-English ASR Dataset
A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.kupe-asr-en-data
kupe-asr-en-mini-150m — data
Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly):
raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this.
mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this.
Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state.
from datasets import load_dataset
ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train")
parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.MonsoonASR-Open-ASR-leaderboard-en-IN
Voice Arena Monsoon en-IN (public test)
Part of the Open ASR Leaderboard, in the main board's default column set, so it contributes to the headline Average WER for every model listed.
A conversational Indian English ASR test set that records who was speaking, not only what
was said. Every clip carries twelve speaker attributes — gender, age, native district
and state, education, occupation, income band, handset — so a difference between two
systems can be traced to a group of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN.ghana-english-speech-ipa
Ghanaian English Speech — Audio with IPA Transcripts
Speech with both transcript forms: the original orthography and the IPA phoneme
sequence read off the audio by ASR. Each language is a subset, with real
train/validation splits.
from datasets import load_dataset
ds = load_dataset("ghanaopendata/ghana-english-speech-ipa", "English_eng", split="train")
ds[0]["audio"] # decoded waveform, 16 kHz
ds[0]["text"] # original orthography
ds[0]["ipa"] # IPA phonemes
52,855… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-ipa.ghana-named-entities-tts-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Named Entities TTS — Twi
A Twi-language speech dataset built from descriptions of Ghana named entities
(people, places, organisations, and concepts). Each audio clip is a synthesised
reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.emilia-en-snac
Stats (EN)
Emilia: 46,349 hours
Emilia-YODAS: 87,258 hours
Total: 133,607 hours
License
The Emilia subset is licensed under CC BY-NC 4.0.
The Emilia-YODAS subset is licensed under CC BY 4.0.
Reference
@inproceedings{emilialarge,
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Wang, Yuancheng and Chen, Kai and Zhang, Pengyuan and Wu… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/emilia-en-snac.enhanced-audiosnippets-long-2-8M
Enhanced Audiosnippets Long 2.8M
Enhanced version of mitermix/audiosnippets_long_2_8M with speech enhancement, emotion annotations, speaker embeddings, and comprehensive metadata analysis.
Dataset Summary
Metric
Value
Total samples
2,633,037
Total audio hours
4,932 h
Duration range
3.0s - 1124.3s
Mean duration
6.7s
Audio format
WAV, 48kHz mono
Tar files
1,410
Processing Pipeline
Each audio sample was processed through:
Speech… See the full description on the dataset page: https://huggingface.co/datasets/ai-music4you3/enhanced-audiosnippets-long-2-8M.mls-eng-speaker-descriptions
Dataset Card for Annotations of English MLS
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.wiktionary-ipa-audio-en
English Wiktionary IPA + audio
English pronunciation rows extracted from the structured Kaikki/Wiktextract
English dump, restricted to English entries with both IPA and a playable
Wikimedia Commons recording. The dataset contains one row per pronunciation
and recording pairing; an audio recording can therefore occur in more than
one row when Wiktionary associates it with multiple IPA or entry records.
Fields
The audio column is created by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/wiktionary-ipa-audio-en.NPTEL_TECH_ENG_50DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translateDisclaimer: The original dataset can be found here.
It is published by Digital Divide Data Cambodia (DDD-Cambodia).
License:
Khmer ASR Cultural Dataset's license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0).
Please attribute Digital Divide Data if you use this dataset in any way.
Objective of this dataset
Add English translation: a new column en_translate is added to the original dataset (only from parquet 000 to 159 of the original… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translate.omnivoice-th
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,833
Total
19,833
entity-transcription-benchmark
Entity Transcription Benchmark
Measures whether a speech recognition system transcribes named entities
correctly — as distinct from word error rate.
WER weights every token equally. The tokens that matter for redaction, lookup,
routing and search are proper nouns, and they are a small fraction of any
transcript. A system can improve WER while getting worse at exactly the words a
downstream consumer needs, and nothing in the standard evaluation will show it.
2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.cv-corpus-1.0-en-client_id-grouped
cv-corpus-1.0-en-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.Meta_STT_EN_Set2
Meta Speech Recognition English Dataset (Set 2)
This dataset contains both metadata and audio files for English speech recognition samples.
Dataset Statistics
Splits and Sample Counts
train: 42961 samples
valid: 2387 samples
test: 2387 samples
Example Samples
train
{
"audio_filepath": "/external1/datasets/asr-himanshu/avspeech-data/audio/AzSutepklXI_2.wav",
"text": "To Jesus, so God is faithful, because when he keeps, you know, when… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN_Set2.omnivoice-zh
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,946
Total
19,946
omnivoice-it
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
CoVoST2-EN-AR
Dataset Description
CoVoST 2 is a large-scale multilingual speech translation corpus based on Common Voice, developed by FAIR. This is the English-to-Arabic portion of the dataset. The original dataset can be found here.
Data Splits (EN-AR)
lang
train
validation
test
EN-AR
289430
15531
15531
AR-EN
2283
1758
1695
Citation
@misc{wang2020covost,
title={CoVoST 2: A Massively Multilingual Speech-to-Text Translation Corpus},
author={Changhan… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/CoVoST2-EN-AR.omnivoice-tr
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
arabic-english-code-switching
Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨
The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning.
Citation
If you use this dataset, please cite it as follows:
@misc{rashad2024arabic,
author = {Mohamed Rashad},
title = {arabic-english-code-switching},
year = {2024},
publisher = {Hugging Face},
url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.disfluency_speech_english
Nyra Disfluency Speech English
nyrahealth/disfluency_speech_english is an English speech dataset for evaluating verbatim ASR: models that should transcribe not only the intended words, but also fillers, cutoffs, repetitions, and sound events.
This dataset is based on the AMAAI Lab DisfluencySpeech dataset and reformatted for verbatim-transcription benchmarking with paired:
verbatim_transcript: what the speaker actually said
intended_transcript: a cleaned version of what the… See the full description on the dataset page: https://huggingface.co/datasets/nyralabs/disfluency_speech_english.mls-eng-128kb
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/ntt123/mls-eng-128kb.omnivoice-es
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
mls-enhanced-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.omnivoice-fr
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
