datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,734
Published sections
493,396
Audio hours
132,613.0
Audio languages
86
Quarantined books
603
Last updated (UTC)
2026-09-24T13:43:06.550876Z
Audio by language
Language
Hours
English
131,663.6
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.tadabur
Tadabur: A Large-Scale Quran Audio Dataset
The most comprehensive and richly annotated Qur'anic recitation corpus to date
Faisal Alherran
✦ Overview
Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur.quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.t2a-daddy
t2a-daddy
Male-voice ASMR corpus for the text2asmr project.
Previously published as aoxo/audios3.
Companion repos: aoxo/t2a-mommy (female voice),
aoxo/t2a-audios-v1 (the original v1 corpus).
Layout
path
what
<creator>/<title>.m4a
source audio, 48 kHz AAC, one folder per creator
<creator>/<title>.json
word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable)
labels/qwen3omni.jsonl
non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-daddy.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.TIE_shorts
Dataset Card for TIE_Shorts
Dataset Summary
TIE_shorts is a derived version of the Technical Indian English (TIE) dataset, a large-scale speech dataset (~ 8K hours) originally consisting of approximately 750 GB of content
sourced from the NPTEL platform. The original TIE dataset contains around 9.8K technical lectures in English delivered by instructors from various regions across India,
with each lecture averaging about 50 minutes. These lectures cover a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/raianand/TIE_shorts.thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.quran-tajweed-phonetics
The complete phonetic layer of the Quran in the riwaya of Hafs 'an
'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every
phone carrying its tajweed attribution: madd class with its transmitted
length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt,
the seventeen sifat, and the rule that produced it.
Built and maintained by Quran Lab, a waqf building open technology in
the service of the Quran.
How it was built and verified
Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.Real-TurnTurk
Real-TurnTurk
English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.mls-eng-speaker-descriptions
Dataset Card for Annotations of English MLS
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.librivox-tracks-silero-vad
librivox-tracks-vad
This dataset is produced from pykeio/librivox-tracks with single-reader filtering and Silero VAD segmentation.
data/train-*.parquet: all utterances (split is always train), collected until a global total audio budget is reached (see run manifest / script args: (2442/5994)*3600 * multiplier seconds by default).
Each row stores source metadata plus a mono WAV payload (audio_bytes) and sampling_rate.
trn-indcnfr-hi-pilot
trn-indcnfr-hi-pilot — Teacher-Output Cache for Ternary-ASR Distillation
Append-only teacher-output cache used to distill a ternary Hindi ASR student.
It stores, per audio clip, ONLY:
row_id (a deterministic source-shard/position pointer),
top-k (k=64) CTC logits (vocab id + log-prob per kept entry, blank always kept),
selected encoder hidden states (last-3 blocks: layers 14 / 15 / 16), fp16.
No audio and no ground-truth transcripts are stored or redistributed.… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/trn-indcnfr-hi-pilot.multilingual-tts-benchmark
Multilingual Speech Benchmark for Zero-Shot TTS
A voice-cloning and intelligibility benchmark for six language variants, built
from Common Voice 17.0 by coverage-driven selection rather than random sampling.
Every example pairs a reference clip of one speaker with a target text that
speaker never read, so a system is asked to clone a voice and produce new
speech, which is what zero-shot TTS is actually for.
Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.tartanaviation-atc-labels
TartanAviation ATC ASR Labels
Machine transcripts and confidence scores for
twangodev/tartanaviation-atc-adsb-utterances.
A 1:1 labels-only add-on (no audio): one row per source utterance, same shards and row order, keyed
by utterance_id.
531,050 labels · 184 shards · 100% coverage · ensemble ASR + weighted ROVER + ADS-B
callsign snap · ~326 human-reviewed. Built with readback.
Usage
Join 1:1 onto the source. Rows are aligned and in the same order:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-labels.quran-audio-text
QuranLab — Verse-Aligned Quran Text + Recitation References
This dataset joins QuranLab's canonical Hafs Arabic text to its
per-ayah recitation references. Every row is one exact
(recitation_id, verse_key) pair: the Uthmani transcript, a search-friendly
Simple-Clean transcript, and the corresponding audio_url.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio-text.bengali-talkshow-audio
Bengali Talkshow Audio Dataset
A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs.
Dataset Description
This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.librivox-tracks-vad
librivox-tracks-vad
This dataset is produced from pykeio/librivox-tracks with single-reader filtering and Silero VAD segmentation.
data/train-*.parquet: all utterances (split is always train), collected until a global total audio budget is reached (see run manifest / script args: (2442/5994)*3600 * multiplier seconds by default).
Each row stores source metadata plus a mono WAV payload (audio_bytes) and sampling_rate.
tadabur
Tadabur: A Large-Scale Quran Audio Dataset
The most comprehensive and richly annotated Qur'anic recitation corpus to date
Faisal Alherran
✦ Overview
Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/MShakir7137/tadabur.transcription-corpus
UN Transcription Corpus
Two splits of UN meeting audio paired with official verbatim records.
Splits
sessions — Whole meeting sessions (SC + GA plenary)
One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org.
Column
Description
symbol
UN document symbol, e.g. S/PV.9826
webtv_url
URL on UN Web TV
duration_ms
Session duration in milliseconds
num_speakers
Number of speaker turns in the verbatim record
audio_floor
Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.arabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.sqp-tts-en
SQP TTS (English)
Synthesized speech for SQPsychConv_qwen-2.5, a synthetic CBT therapist-client
dialogue dataset (English).
Each configuration below corresponds to one TTS model. Load a single model
with:
from datasets import load_dataset
ds = load_dataset("sinselm/sqp-tts-en", "qwen3-tts")
Models included
qwen3-tts: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base
cosyvoice: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512
fishaudio:… See the full description on the dataset page: https://huggingface.co/datasets/marleen-snsl/sqp-tts-en.script-fidelity-benchmark
Script fidelity benchmark
Anonymous supplement for the paper "Script collapse in multilingual ASR:
A reference-free metric and 100-pair benchmark."
Script Fidelity Rate (SFR) measures the fraction of ASR hypothesis characters
that belong to the expected target script. WER measures word edits, while SFR
checks whether the output is written in the target orthography.
Related resources:
PyPI package: https://pypi.org/project/script-fidelity/
Hugging Face Evaluate metric:… See the full description on the dataset page: https://huggingface.co/datasets/themechanism/script-fidelity-benchmark.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.sun-yuchen-selected-works
孙宇晨文选音频文本语料库
这是一个由 157 条中文长音频、157 篇 AI 清洗 Markdown 和 Whisper 毫秒级时间戳整理而成的 Hugging Face 音频数据集仓库。它同时提供长节目级 episodes 和时间戳片段级 segments 两个配置,用于选集制作、语料分析、自动语音识别(ASR)研究,以及经过额外人工审核后的文本转语音(TTS)数据准备。
发布边界: 本数据集仓库可公开访问,但当前源音频的许可、再分发授权、来源平台条款及声音/人格权尚未确认。仓库使用 license: other,公开访问不等于授予任何数据权利;未经权利核验与授权,不应商用、再分发或用于声音克隆。详见 LICENSE。
数据集摘要
配置
样本单位
train
validation
test
总计
时长
episodes
完整节目
125
16
16
157
31.234 小时
segments
时间戳语音片段
7,738
864
495
9,097
30.932 小时… See the full description on the dataset page: https://huggingface.co/datasets/tradecatlabs/sun-yuchen-selected-works.egyptian-arabic-tts-diacritized
Egyptian Arabic TTS Corpus (Diacritized)
97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with
diacritized transcripts — the short vowels that Arabic script does not
write.
Why diacritics
Arabic is an abjad: short vowels are unwritten, so كتب may be kataba,
kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which,
and guesses — which native listeners hear as a foreign accent with constant
mispronunciation.
This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-10k-tags_tagged_10k_generated.massive-yt-edu-queue
Massive YouTube Educational Video Queue
Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours.
Description
This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.
