datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,734
Published sections
493,396
Audio hours
132,613.0
Audio languages
86
Quarantined books
603
Last updated (UTC)
2026-09-24T13:43:06.550876Z
Audio by language
Language
Hours
English
131,663.6
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.LEMAS-Dataset-train
Overview
This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.TAGARELA
TAGARELA: A Portuguese Speech Dataset From Podcasts
TAGARELA is a large-scale Portuguese speech dataset built from podcast audio and curated for speech technology research, especially Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
The dataset contains more than 8,972 hours of Portuguese speech derived from the Cem Mil Podcasts collection. It includes Brazilian Portuguese and European Portuguese speech, processed through a pipeline involving audio standardization… See the full description on the dataset page: https://huggingface.co/datasets/freds0/TAGARELA.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.tat_youtubeopen-yap-1k
Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use
Today we're releasing Open Yap 1K: 1,000 hours of dual-channel English conversation, capturing how people speak together naturally in real-world environments recorded in 48kHz. The dataset ships free for both commercial and research use.
The sample on the Hugging Face Hub - 8.9 hours, 16 conversations, CC-BY-4.0, listenable in the dataset viewer.
The full corpus - 1,000 hours, 1,602… See the full description on the dataset page: https://huggingface.co/datasets/TheAgenticDataCompany/open-yap-1k.libritts_r_filtered
Dataset Card for Filtered LibriTTS-R
This is a filtered version of LibriTTS-R. It has been filtered based on two sources:
LibriTTS-R paper [1], which lists samples for which speech restoration have failed
LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected.
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately
585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.ramanv-tts-all-raw
ramanv-tts-all-raw
Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages.
police-scanner-audio
Police Scanner Audio Dataset
A comprehensive collection of police and emergency services radio communications from multiple US cities, captured from publicly available scanner feeds.
Dataset Overview
This dataset contains 103,660 audio recordings totaling 357GB of police scanner audio from 6 different cities across the United States. The recordings span multiple months of continuous monitoring and represent real-world emergency services communications.
Scanner… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/police-scanner-audio.tadabur
Tadabur: A Large-Scale Quran Audio Dataset
The most comprehensive and richly annotated Qur'anic recitation corpus to date
Faisal Alherran
✦ Overview
Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur.everyayah﷽
Dataset Card for Tarteel AI's EveryAyah Dataset
Dataset Summary
This dataset is a collection of Quranic verses and their transcriptions, with diacritization, by different reciters.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in Arabic.
Dataset Structure
Data Instances
A typical data point comprises the audio file audio, and its transcription called text.
The duration… See the full description on the dataset page: https://huggingface.co/datasets/tarteel-ai/everyayah.emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.kasem-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Kasem Speech-Text Parallel Dataset
Dataset Description
This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.t2a-daddy
t2a-daddy
Male-voice ASMR corpus for the text2asmr project.
Previously published as aoxo/audios3.
Companion repos: aoxo/t2a-mommy (female voice),
aoxo/t2a-audios-v1 (the original v1 corpus).
Layout
path
what
<creator>/<title>.m4a
source audio, 48 kHz AAC, one folder per creator
<creator>/<title>.json
word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable)
labels/qwen3omni.jsonl
non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-daddy.open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.linustechtips-transcript-audio
Dataset Card for "linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.FalaBracarense_splitsdataset website: projectofalabracarense
Licence
CC - BY - NC - ND
Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives
tts_farm
Multilingual TTS/ASR Aggregated Dataset
Cleaned, deduplicated and loudness-normalized Arabic, Japanese, Korean, Turkish, and Vietnamese speech. The training columns are audio (16-bit PCM WAV, 22050 Hz) and text; the remaining columns contain quality and provenance metadata.
UltraVoice
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
📝 Abstract
Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question answering. To address this limitation, we introduce UltraVoice, the first large-scale speech dialogue dataset… See the full description on the dataset page: https://huggingface.co/datasets/tutu0604/UltraVoice.majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.lex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.Persian-Farsi-Speech
Persian (Farsi) TTS Dataset
🗂️ Dataset Description
This dataset is a Persian (Farsi) text-to-speech (TTS) corpus built by concatenating and denoising multiple existing Farsi datasets.It is intended for training and evaluation of speech synthesis (TTS) models in Persian.
Since the basic datasets were contaminated with unintelligible audio, I used dnsmos to keep only clean audio (mos_ovr >= 3.0, same value as for the Emilia dataset).
The dataset contains two main… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/Persian-Farsi-Speech.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/adab-tech/WaxalNLP.Vaani-transcription-partThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages.
This table represents the audio and transcription duration data for various languages.
Language
Angami
Angika
Ao
Assamese
Awadhi
Bajjika
Bearybashe
Bengali
Bhili
Bhojpuri
Bundeli
Chakhesang
Chakma
Chhattisgarhi
English
Garhwali
Garo
Gondi
Gujarati
Halbi
Haryanvi
Hindi
IduMishmi
Kannada
Kashmiri
Karbi
Khariboli
Khortha
Kokborok
Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.WenetSpeech-Formal
WenetSpeech-Formal
Dataset Description
WenetSpeech-Formal is a Chinese speech dataset built upon WenetSpeech, augmented with formal written-style transcriptions. Each audio sample contains two levels of transcription:
Field
Description
original_text
Original verbatim transcription(口语原文,含口语习惯表达)
target_text
Formal written-style transcription(正式书面体转写,由大语言模型辅助生成)
This dataset is designed for training models that convert spoken Chinese to formal… See the full description on the dataset page: https://huggingface.co/datasets/TaurenMountain/WenetSpeech-Formal.
