datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MIT_environmental_impulse_responsesMIT Environmental Impulse Response Dataset
The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html.
The audio files in the dataset have been resampled to a sampling rate of 16 kHz. This resampling was done to reduce the size of the dataset while making it more suitable for various tasks, including data augmentation.
The dataset consists of 271 audio files… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/MIT_environmental_impulse_responses.mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.ghana-english-asr-2700hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.ghana-english-speech-600hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-600hrs.translated-german-english-asr
Translated German-English ASR Dataset
A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.MonsoonASR-Open-ASR-leaderboard-en-IN
Voice Arena Monsoon en-IN (public test)
Part of the Open ASR Leaderboard, in the main board's default column set, so it contributes to the headline Average WER for every model listed.
A conversational Indian English ASR test set that records who was speaking, not only what
was said. Every clip carries twelve speaker attributes — gender, age, native district
and state, education, occupation, income band, handset — so a difference between two
systems can be traced to a group of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN.ghana-english-speech-ipa
Ghanaian English Speech — Audio with IPA Transcripts
Speech with both transcript forms: the original orthography and the IPA phoneme
sequence read off the audio by ASR. Each language is a subset, with real
train/validation splits.
from datasets import load_dataset
ds = load_dataset("ghanaopendata/ghana-english-speech-ipa", "English_eng", split="train")
ds[0]["audio"] # decoded waveform, 16 kHz
ds[0]["text"] # original orthography
ds[0]["ipa"] # IPA phonemes
52,855… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-ipa.ghana-named-entities-tts-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Named Entities TTS — Twi
A Twi-language speech dataset built from descriptions of Ghana named entities
(people, places, organisations, and concepts). Each audio clip is a synthesised
reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.MIT_environmental_impulse_responses
MIT Environmental Impulse Response Dataset
The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html.
This mirror provides the 16 kHz WAV files used for wake-word training augmentation in the Tater Totterson trainer projects. The files were resampled to 16 kHz to keep the dataset small and convenient for machine-learning audio pipelines.… See the full description on the dataset page: https://huggingface.co/datasets/TaterTotterson/MIT_environmental_impulse_responses.Datasets_ENwiktionary-ipa-audio-en
English Wiktionary IPA + audio
English pronunciation rows extracted from the structured Kaikki/Wiktextract
English dump, restricted to English entries with both IPA and a playable
Wikimedia Commons recording. The dataset contains one row per pronunciation
and recording pairing; an audio recording can therefore occur in more than
one row when Wiktionary associates it with multiple IPA or entry records.
Fields
The audio column is created by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/wiktionary-ipa-audio-en.DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translateDisclaimer: The original dataset can be found here.
It is published by Digital Divide Data Cambodia (DDD-Cambodia).
License:
Khmer ASR Cultural Dataset's license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0).
Please attribute Digital Divide Data if you use this dataset in any way.
Objective of this dataset
Add English translation: a new column en_translate is added to the original dataset (only from parquet 000 to 159 of the original… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translate.omnivoice-th
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,833
Total
19,833
entity-transcription-benchmark
Entity Transcription Benchmark
Measures whether a speech recognition system transcribes named entities
correctly — as distinct from word error rate.
WER weights every token equally. The tokens that matter for redaction, lookup,
routing and search are proper nouns, and they are a small fraction of any
transcript. A system can improve WER while getting worse at exactly the words a
downstream consumer needs, and nothing in the standard evaluation will show it.
2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.cv-corpus-1.0-en-client_id-grouped
cv-corpus-1.0-en-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.Meta_STT_EN_Set2
Meta Speech Recognition English Dataset (Set 2)
This dataset contains both metadata and audio files for English speech recognition samples.
Dataset Statistics
Splits and Sample Counts
train: 42961 samples
valid: 2387 samples
test: 2387 samples
Example Samples
train
{
"audio_filepath": "/external1/datasets/asr-himanshu/avspeech-data/audio/AzSutepklXI_2.wav",
"text": "To Jesus, so God is faithful, because when he keeps, you know, when… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN_Set2.omnivoice-zh
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,946
Total
19,946
omnivoice-it
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
CoVoST2-EN-AR
Dataset Description
CoVoST 2 is a large-scale multilingual speech translation corpus based on Common Voice, developed by FAIR. This is the English-to-Arabic portion of the dataset. The original dataset can be found here.
Data Splits (EN-AR)
lang
train
validation
test
EN-AR
289430
15531
15531
AR-EN
2283
1758
1695
Citation
@misc{wang2020covost,
title={CoVoST 2: A Massively Multilingual Speech-to-Text Translation Corpus},
author={Changhan… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/CoVoST2-EN-AR.omnivoice-tr
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
arabic-english-code-switching
Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨
The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning.
Citation
If you use this dataset, please cite it as follows:
@misc{rashad2024arabic,
author = {Mohamed Rashad},
title = {arabic-english-code-switching},
year = {2024},
publisher = {Hugging Face},
url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.disfluency_speech_english
Nyra Disfluency Speech English
nyrahealth/disfluency_speech_english is an English speech dataset for evaluating verbatim ASR: models that should transcribe not only the intended words, but also fillers, cutoffs, repetitions, and sound events.
This dataset is based on the AMAAI Lab DisfluencySpeech dataset and reformatted for verbatim-transcription benchmarking with paired:
verbatim_transcript: what the speaker actually said
intended_transcript: a cleaned version of what the… See the full description on the dataset page: https://huggingface.co/datasets/nyralabs/disfluency_speech_english.mls-eng-128kb
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/ntt123/mls-eng-128kb.omnivoice-es
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
mls-enhanced-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.omnivoice-fr
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
sqp-tts-en
SQP TTS (English)
Synthesized speech for SQPsychConv_qwen-2.5, a synthetic CBT therapist-client
dialogue dataset (English).
Each configuration below corresponds to one TTS model. Load a single model
with:
from datasets import load_dataset
ds = load_dataset("sinselm/sqp-tts-en", "qwen3-tts")
Models included
qwen3-tts: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base
cosyvoice: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512
fishaudio:… See the full description on the dataset page: https://huggingface.co/datasets/marleen-snsl/sqp-tts-en.Ghana_English-Twi_Code-switching_Speech
Dataset Card for KasaSpeech
Dataset Summary
KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi.
The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios
With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code-switching_Speech.omnivoice-ja
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,958
Total
19,958
