datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yodas2_sidon
YODAS2-Sidon
Overview
This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks.
We resampled original sidon output to 24kHz due to a storage constraints.
The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.rixvox-v2
RixVox-v2: A Swedish parliamentary speech dataset
RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.gigaspeech2
Dataset Card for GigaSpeech 2
Dataset Description
GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese.
Repository: https://github.com/SpeechColab/GigaSpeech2
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.cv_corpus_v22
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
NOTE: currently converting to parquet for convenience.. WIP
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.japanese-anime-speech-v2
Japanese Anime Speech Dataset V2
日本語はこちら
japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models.
The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels.
This dataset is not an updated version of japanese-anime-speech-v1.
For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset.
The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2.WorldSpeech
WorldSpeech
A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR estimate, and four DNSMOS-P.835 quality scores.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Centi234/WorldSpeech.dialectal-arabic-lahgtna-v2
Dialectal Arabic Lahgtna v2
Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI.
Dataset Summary
~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech
**13 Arabic dialects **, labeled per utterance
16 kHz mono audio
Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.StreamAudio-2M
StreamAudio-2M
Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a
stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips
are organised into six task subsets.
Subsets
Subset
Rows
Description
Stream_Audio_Understanding
90,738
Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA
Real_time_ASR
28,109
Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.Voices-in-the-Wild-2M
Voices in the Wild
Project Page | Paper | GitHub
Voices in the Wild (Voices-in-the-Wild-2M) is a large-scale automatic speech recognition (ASR) dataset designed for robustness training and evaluation under diverse, real-world acoustic conditions. It covers 7 classic acoustic phenomena (including noise, far-field speech, obstruction, echo/reverberation, recording artifacts, electronic distortion, and transmission dropout) and 54 physically plausible compound scenarios.
The… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/Voices-in-the-Wild-2M.GLOBE_V2
Important notice
Differences between V2 version and the version described in paper:
The V2 version provide audio in 44.1kHz sample rate. (Supersampling)
The V2 versionn removed some samples (~5%) due to the volumn and text aligment issues.
Globe
The full paper can be accessed here: arXiv
An online demo can be accessed here: Github
Abstract
This paper introduces GLOBE, a high-quality English corpus with worldwide accents, specifically designed to address the… See the full description on the dataset page: https://huggingface.co/datasets/MushanW/GLOBE_V2.ghana-english-asr-2700hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.yt-danish-public-v22M-Belebele
2M-Belebele
Highly-Multilingual Speech and American Sign Language Comprehension Dataset
We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL).
The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.SADA22
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of transcribed Arabic audio recordings, primarily featuring various Saudi dialects, and was curated in a collaboration between the National Center for Artificial… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/SADA22.SPGISpeech2.0
Dataset Card for SPGISpeech 2.0
Dataset Details
Dataset Overview
We are excited to present SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR).… See the full description on the dataset page: https://huggingface.co/datasets/kensho/SPGISpeech2.0.Kazakh_Speech_Corpus_2
Kazakh Speech Corpus 2 (KSC2)
This dataset card describes the KSC2, an industrial-scale, open-source speech corpus for the Kazakh language.
Paper: KSC2: An Industrial-Scale Open-Source Kazakh Speech Corpus
Summary: KSC2 corpus subsumes the previously introduced two corpora: Kazakh Speech Corpus and Kazakh Text-To-Speech 2, and supplements additional data from other sources like tv programs, radio, senate, and podcasts. In total, KSC2 contains around 1.2k hours of high-quality… See the full description on the dataset page: https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2.EuroSpeech-24kHz
EuroSpeech 24 kHz Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper.
Dataset Summary
Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.gemini-flash-2.0-speech
🎙️ Gemini Flash 2.0 Speech Dataset
This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English.
🏅 #1 Trending Audio Dataset in Feb 2025
🏅 Used in training of Kokoro TTS and LLaSA 1B
〽️ Stats
Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours)
Average duration: 10.83 seconds
Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.mgb2-arabic
MGB-2: Arabic Multi-Dialect Broadcast Media Recognition
Dataset Description
Dataset Summary
The Arabic Multi-Genre Broadcast (MGB-2) dataset is a large-scale speech recognition corpus containing 1,200 hours of Arabic broadcast audio from Aljazeera Arabic TV channel. The dataset spans recordings from March 2005 to December 2015 and covers 19 distinct programme series. It was originally created for the MGB-2 Challenge at SLT-2016, focusing on handling dialect… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/mgb2-arabic.afri-temp-data4
AfricanVoices Hausa -- Train Split
Hausa speech dataset from AfricanVoices.io.
Usage
from datasets import load_dataset
ds = load_dataset("suleiman2003/afri-temp-data4", split="train")
print(ds[0])
# {'audio': Audio(...), 'transcript': '...', 'gender': '...', ...}
Structure
Each batch is in its own subdirectory under train/:
train/
batch_1/
*.flac + metadata.csv
batch_2/
*.flac + metadata.csv
...
Audio files are FLAC format. Metadata… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/afri-temp-data4.earnings25
Earnings25
A 500-hour speech benchmark for finance — S&P 500 earnings calls with reference
transcripts, industry labels, and named-speaker attribution.
Citation
Earnings25 is introduced in our Interspeech 2026 paper,
which sets out the sampling design, the evaluation protocol, and reference
baselines for Whisper and Parakeet-TDT. Start there for the full picture.
Jiang, D., Zhou, H., Wadhawan, A., Fahy, B., Ramesh, V., Weisberg, D.,
Derkachevskiy, D., Sheehan, H., Prasad, S., &… See the full description on the dataset page: https://huggingface.co/datasets/florencejiang/earnings25.commonvoice22-sidon-dacvae
CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents
Source
sarulab-speech/commonvoice22_sidon
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.DNU-ViSpeech
DNU-ViSpeech
Publication status: The manuscript describing DNU-ViSpeech is currently under review and has not yet been accepted for publication. This repository provides the dataset and associated benchmark results for research and reproducibility. The paper citation will be added when publicly available.
Public release: the dataset owner authorized public release on 2026-09-05 after attesting that all applicable human-review gates were completed. See… See the full description on the dataset page: https://huggingface.co/datasets/TrangLe2108/DNU-ViSpeech.Quran-Ayah-Corpus
Quran-Ayah-Corpus: A Multi-Reciter Arabic Quranic Speech Dataset
Dataset Description:
Ayah-Corpus is a large-scale, multi-reciter Arabic speech dataset meticulously curated for Automatic Speech Recognition (ASR) tasks. It consists of high-quality audio recordings of Quranic verses (Ayahs) paired with their corresponding exact transcriptions. The audio is sourced from two primary repositories: Al-Quran.cloud and EveryAyah.com.
This dataset is specifically designed to… See the full description on the dataset page: https://huggingface.co/datasets/rabah2026/Quran-Ayah-Corpus.infore2_audiobooks
unofficial mirror of InfoRe Technology public dataset №2
official announcement: https://www.facebook.com/groups/j2team.community/permalink/1010834009248719/
415h, 315k samples, vietnamese audiobooks of chinese wǔxiá 武俠 & xiānxiá 仙俠
bộ dữ liệu bóc ra từ YouTube đọc truyện võ hiệp & tiên hiệp, áp dụng kĩ thuật đối chiếu văn bản để dán nhãn tự động
official download:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/infore2_audiobooks.s2m-bang
s2m-bang
Bangla Stage-1 speech→transcript dataset used for speech-to-LLM adapter training
(Moonshine-BN encoder + LLM adapter).
Splits
Manifest
n
Role
manifests/train.json
89,275
Train (Kathbath + CV/OpenSLR short + FLEURS train)
manifests/dev.json
2,638
Dev
manifests/dev_fast.json
256
Fast mid-train gate
manifests/fleurs_test.json
920
Held-out FLEURS test
manifests/cv_bn_eval.json
260
Held-out Common Voice eval
Schema
Each… See the full description on the dataset page: https://huggingface.co/datasets/niloy629/s2m-bang.afri-temp-data3
AfricanVoices Hausa -- Train Split
from datasets import load_dataset
ds = load_dataset("suleiman2003/afri-temp-data3", split="train")
print(ds[0])
reazon-speech-v2-clone
Reazon Speech v2 dataset mirror
Original Dataset Source
Hugging Face Dataset Page: reazon-research/reazonspeech
Project Page: Reazon Research
License
This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction:
TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.
