datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,728
Published sections
493,262
Audio hours
132,580.7
Audio languages
86
Quarantined books
606
Last updated (UTC)
2026-09-23T13:55:55.728817Z
Audio by language
Language
Hours
English
131,631.3
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.DeepDialogue-orpheus
DeepDialogue-orpheus
DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text.
🚨 Important Notice
This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.tadabur
Tadabur: A Large-Scale Quran Audio Dataset
The most comprehensive and richly annotated Qur'anic recitation corpus to date
Faisal Alherran
✦ Overview
Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur.quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.ytseg
YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation
We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.Bagpiper_PreTrain_Data
Bagpiper Pretraining Data
Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated
with Bagpiper, an open-ended audio language
model that learns bidirectional mappings between audio and comprehensive text
descriptions across speech, music, environmental sound, and mixtures.
The en metadata describes the primary rich-caption language. Source audio can
contain speech or singing in other languages; it is not an English-only audio
guarantee.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.ParsVoice
ParsVoice
A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
📣 Accepted to the EMNLP 2026 Main Conference.
ParsVoice is the largest publicly available Persian speech–text corpus tailored for
training multi-speaker text-to-speech (TTS) systems. It is built from long-form
Persian audiobook recordings using a fully automated pipeline combining sentence-aware
segmentation, ASR transcription, a ParsBERT sentence-completion classifier, binary-search… See the full description on the dataset page: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice.2M-Belebele
2M-Belebele
Highly-Multilingual Speech and American Sign Language Comprehension Dataset
We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL).
The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.danish-asr-leaderboard
Open Danish ASR Leaderboard — Results
Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models.
Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.risale-i-nur-sohbet
Risale-i Nur Sohbet
Prof. Dr. Şener Dilek’ten izin alındı.
Türkçe
Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde
birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış
sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri
kümelerine karıştırılmaz.
Kapsam
2095 sohbet, 954.66 saat 16 kHz mono FLAC ses
Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste
seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.kupe-asr-en-data
kupe-asr-en-mini-150m — data
Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly):
raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this.
mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this.
Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state.
from datasets import load_dataset
ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train")
hashed_data
Munch Hashed Index - Lightweight Audio Reference Dataset
📖 Overview
Munch Hashed Index is a lightweight reference dataset that provides SHA-256 hashes for all audio files in the Munch Urdu TTS Dataset. Instead of storing 1.27 TB of raw audio, this index stores only metadata and cryptographic hashes, enabling:
✅ Fast duplicate detection across 4.17 million audio samples
✅ Efficient dataset exploration without downloading terabytes
✅ Quick metadata queries (voice… See the full description on the dataset page: https://huggingface.co/datasets/humair025/hashed_data.MyMentorLLM-dataset
Dataset Card for MyMentorLLM
This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information).
Dataset Summary
MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.TIE_shorts
Dataset Card for TIE_Shorts
Dataset Summary
TIE_shorts is a derived version of the Technical Indian English (TIE) dataset, a large-scale speech dataset (~ 8K hours) originally consisting of approximately 750 GB of content
sourced from the NPTEL platform. The original TIE dataset contains around 9.8K technical lectures in English delivered by instructors from various regions across India,
with each lecture averaging about 50 minutes. These lectures cover a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/raianand/TIE_shorts.seamless-interaction-jefferson-annotations
Seamless Interaction Jefferson-Style Annotations
An automatic, turn-oriented annotation layer for the
Meta Seamless Interaction Dataset.
It compares the dataset's traditional transcript with an ASR-derived
Jefferson-style condition and supplies speech-act, communicative-purpose,
interactional-signal, alignment, and quality fields.
This is a derived noncommercial research dataset. It does not redistribute
the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.YouTube-Cantonese-Emilia
YouTube Cantonese — Emilia
2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by
running alvanlii/cantonese-youtube
through the Emilia
speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering).
Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn
label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in
both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.hk-legicost
HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation
HK-LegiCoST is a three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and aligned at the sentence level.
Paper: arXiv:2306.11252
Authors: Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao, Matthew Wiesner, Kevin Duh, Sanjeev Khudanpur
Dataset Description
The raw… See the full description on the dataset page: https://huggingface.co/datasets/Borrison/hk-legicost.Whisper-Hallucination
Whisper Hallucination and Repetition Probes
This is a benchmark. Every evaluation config is test — do not fine-tune on it.
lexicon_synth is the exception: synthetic training material with its own train/test
split, and not one of the eight benchmark arms.
To build training data, exclude the items in
benchmark/exclusions.json
(546 FMA tracks, 1,168 FSD50K ids, 2,620 LibriSpeech utterances, the Malay stems). The
benchmark draws FSD50K eval and FMA shards 0–1, so training can use… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.omniasr-molge
OmniASR Molge Aligned
Training-friendly re-segmentation of Meta’s facebook/omnilingual-asr-corpus: long utterances are segmented / aligned into ≤30s clips with transcripts, then packed as Parquet shards with embedded FLAC.
Source
facebook/omnilingual-asr-corpus
Configs
omniasr_aligned_v1, omniasr_aligned_v2
Splits
train / validation (dev-*.parquet) / test
Scale
~2.56M utts · ~839 shards · ~439GB
If this dataset is useful for your work, we’d appreciate a… See the full description on the dataset page: https://huggingface.co/datasets/Sanghyang00/omniasr-molge.yodas3
YODAS v3
Paper
YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 and v2 versions of YODAS, to guarantee that there are no overlaps in the data.
For… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas3.danish-asr-verified
danish-asr-verified
ALL rows of syvai/danish-asr-unified transcribed by the
syv-transcribe ensemble (hviske-v5.3 + hviske-v5, confidence-weighted ROVER),
each annotated with:
verified — True when the ensemble independently reproduced the reference
exactly (compared after lowercasing, punctuation-strip, whitespace-collapse).
Two independent witnesses agree => near-certain label.
wer_teacher_vs_ref / cer_teacher_vs_ref — word/character error rate
between normalized teacher output… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-verified.omnievalkit-dataset
OmniEvalKit Evaluation Datasets
Evaluation datasets for OmniEvalKit,
a comprehensive evaluation framework for omni-modal (audio + video + image + text) models.
Overview
Total subsets: 65
Total samples: 315,264
Total size: 620.3 GB (Parquet with embedded audio/image/video)
Subsets with embedded video: 15
Subsets requiring external video download: 2
Usage
from datasets import load_dataset
ds = load_dataset("OmniEvalKit/omnievalkit-dataset", "aishell1_test")… See the full description on the dataset page: https://huggingface.co/datasets/OmniEvalKit/omnievalkit-dataset.arknights_voices_zh
ZH Voice-Text Dataset for Arknights Waifus
This is the ZH voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
12431 records, 25.9 hours in total. Average duration is approximately 7.49s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_106_franka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_zh.enhanced-audiosnippets-long-2-8M
Enhanced Audiosnippets Long 2.8M
Enhanced version of mitermix/audiosnippets_long_2_8M with speech enhancement, emotion annotations, speaker embeddings, and comprehensive metadata analysis.
Dataset Summary
Metric
Value
Total samples
2,633,037
Total audio hours
4,932 h
Duration range
3.0s - 1124.3s
Mean duration
6.7s
Audio format
WAV, 48kHz mono
Tar files
1,410
Processing Pipeline
Each audio sample was processed through:
Speech… See the full description on the dataset page: https://huggingface.co/datasets/ai-music4you3/enhanced-audiosnippets-long-2-8M.svq
Simple Voice Questions
Simple Voice Questions (SVQ) is a set of short audio questions recorded in 26 locales across 17 languages under multiple audio conditions.
Data Collection
Speakers were presented with recording instructions specifying the recording environment and text query to be recorded.
They recorded using their own phones or tablets under four conditions:
clean: Record in quiet environment
background speech noise: Record while audio from sources like podcasts… See the full description on the dataset page: https://huggingface.co/datasets/mteb/svq.
