datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yodas2_sidon
YODAS2-Sidon
Overview
This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks.
We resampled original sidon output to 24kHz due to a storage constraints.
The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.common_voice_22_0
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_22_0.voxceleb2
VoxCeleb2 Dataset
This is the VoxCeleb2 dataset, a large-scale speaker identification dataset.
Dataset Description
VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube.
Files
vox2_dev_mp4_part*: Multipart archive containing MP4 video files
vox2_dev_txt: Text files with speaker/utterance metadata
vox2_meta.csv: Dataset metadata
Usage
To extract the multipart archive:
# Using 7zip
7z x… See the full description on the dataset page: https://huggingface.co/datasets/Reverb/voxceleb2.rixvox-v2
RixVox-v2: A Swedish parliamentary speech dataset
RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.Thinkspark-v2-270m-training-data
ThinkSpark-v2-350M — training data
Full-duplex floor-controller (Section 8) training corpus: playable audio + text,
paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain,
gender, prosody, agent text) and Soniox character-level timestamps.
Dataset Viewer
Default split is parquet with a real Audio feature — a player renders inline next to
the text in the Hub UI:
column
type
description
audio
Audio
playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.gigaspeech2
Dataset Card for GigaSpeech 2
Dataset Description
GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese.
Repository: https://github.com/SpeechColab/GigaSpeech2
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.cv_corpus_v22
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
NOTE: currently converting to parquet for convenience.. WIP
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.japanese-anime-speech-v2
Japanese Anime Speech Dataset V2
日本語はこちら
japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models.
The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels.
This dataset is not an updated version of japanese-anime-speech-v1.
For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset.
The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2.WorldSpeech
WorldSpeech
A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR estimate, and four DNSMOS-P.835 quality scores.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Centi234/WorldSpeech.dialectal-arabic-lahgtna-v2
Dialectal Arabic Lahgtna v2
Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI.
Dataset Summary
~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech
**13 Arabic dialects **, labeled per utterance
16 kHz mono audio
Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.StreamAudio-2M
StreamAudio-2M
Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a
stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips
are organised into six task subsets.
Subsets
Subset
Rows
Description
Stream_Audio_Understanding
90,738
Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA
Real_time_ASR
28,109
Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.Voices-in-the-Wild-2M
Voices in the Wild
Project Page | Paper | GitHub
Voices in the Wild (Voices-in-the-Wild-2M) is a large-scale automatic speech recognition (ASR) dataset designed for robustness training and evaluation under diverse, real-world acoustic conditions. It covers 7 classic acoustic phenomena (including noise, far-field speech, obstruction, echo/reverberation, recording artifacts, electronic distortion, and transmission dropout) and 54 physically plausible compound scenarios.
The… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/Voices-in-the-Wild-2M.afrispeech-200AFRISPEECH-200 is a 200hr Pan-African speech corpus for clinical and general domain English accented ASR;
a dataset with 120 African accents from 13 countries and 2,463 unique African speakers.
Our goal is to raise awareness for and advance Pan-African English ASR research,
especially for the clinical domain.GLOBE_V2
Important notice
Differences between V2 version and the version described in paper:
The V2 version provide audio in 44.1kHz sample rate. (Supersampling)
The V2 versionn removed some samples (~5%) due to the volumn and text aligment issues.
Globe
The full paper can be accessed here: arXiv
An online demo can be accessed here: Github
Abstract
This paper introduces GLOBE, a high-quality English corpus with worldwide accents, specifically designed to address the… See the full description on the dataset page: https://huggingface.co/datasets/MushanW/GLOBE_V2.ghana-english-asr-2700hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.yt-danish-public-v22M-Belebele
2M-Belebele
Highly-Multilingual Speech and American Sign Language Comprehension Dataset
We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL).
The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.SADA22
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of transcribed Arabic audio recordings, primarily featuring various Saudi dialects, and was curated in a collaboration between the National Center for Artificial… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/SADA22.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.SPGISpeech2.0
Dataset Card for SPGISpeech 2.0
Dataset Details
Dataset Overview
We are excited to present SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR).… See the full description on the dataset page: https://huggingface.co/datasets/kensho/SPGISpeech2.0.Kazakh_Speech_Corpus_2
Kazakh Speech Corpus 2 (KSC2)
This dataset card describes the KSC2, an industrial-scale, open-source speech corpus for the Kazakh language.
Paper: KSC2: An Industrial-Scale Open-Source Kazakh Speech Corpus
Summary: KSC2 corpus subsumes the previously introduced two corpora: Kazakh Speech Corpus and Kazakh Text-To-Speech 2, and supplements additional data from other sources like tv programs, radio, senate, and podcasts. In total, KSC2 contains around 1.2k hours of high-quality… See the full description on the dataset page: https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2.common_voice_21_0
Dataset Card for Common Voice Corpus 21.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 21. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash… See the full description on the dataset page: https://huggingface.co/datasets/Theafricatechguy/common_voice_21_0.audio-v2This dataset contains >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It wa released on April 20th, 2025.
You can find the full list of sources in this dataset under the dataset's sources.txt.
Paper: https://arxiv.org/abs/2307.08720
If you use our datasets, the following quote is preferable:
@misc{marmor2023ivritai,
title={ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development},
author={Yanir Marmor and Kinneret Misgav and Yair… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2.EuroSpeech-24kHz
EuroSpeech 24 kHz Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper.
Dataset Summary
Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.2M-Flores-ASL
2M-Flores
As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest
sentences in the original flores200 dataset.
To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded.
The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time.
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.common_voice_21_0
Dataset Card for Common Voice Corpus 21.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 21. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_21_0.gemini-flash-2.0-speech
🎙️ Gemini Flash 2.0 Speech Dataset
This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English.
🏅 #1 Trending Audio Dataset in Feb 2025
🏅 Used in training of Kokoro TTS and LLaSA 1B
〽️ Stats
Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours)
Average duration: 10.83 seconds
Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.afri-temp-data4
AfricanVoices Hausa -- Train Split
Hausa speech dataset from AfricanVoices.io.
Usage
from datasets import load_dataset
ds = load_dataset("suleiman2003/afri-temp-data4", split="train")
print(ds[0])
# {'audio': Audio(...), 'transcript': '...', 'gender': '...', ...}
Structure
Each batch is in its own subdirectory under train/:
train/
batch_1/
*.flac + metadata.csv
batch_2/
*.flac + metadata.csv
...
Audio files are FLAC format. Metadata… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/afri-temp-data4.
