datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
svq
Simple Voice Questions
Simple Voice Questions (SVQ) is a set of short audio questions recorded in 26 locales across 17 languages under multiple audio conditions. It serves as a core evaluation componenet for Massive Sound Embedding Benchmark (MSEB).
Technical Specifications
Feature
Details
Locales
26
Languages
17
Total Speakers
~700 (Capped at 250 recordings per speaker)
Audio Conditions
Clean, Background Speech, Media, Traffic Noise
Gender… See the full description on the dataset page: https://huggingface.co/datasets/google/svq.librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,724
Published sections
493,186
Audio hours
132,549.7
Audio languages
86
Quarantined books
610
Last updated (UTC)
2026-09-21T15:06:16.273192Z
Audio by language
Language
Hours
English
131,600.3
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.Granary
Granary: Speech Recognition and Translation Dataset in 25 European Languages
Granary is a large-scale, open-source multilingual speech dataset covering 25 European languages for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) tasks.
Overview
Granary addresses the scarcity of high-quality speech data for low-resource languages by consolidating multiple datasets under a unified framework:
🗣️ ~1M hours of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Granary.Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.QuranTTS
QuranTTS v4
An ear-verified Quranic recitation corpus for TTS and speech restoration.
Built in-house for the lab's own training runs. We have since moved on to a larger corpus and a newer
pipeline, so this one is published rather than shelved. What you get is the corpus exactly as it stood when
we stopped using it: complete, documented, and unmaintained.
Non-commercial, strictly. Neither this corpus nor any model trained on it may be used for any
commercial purpose, and… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/QuranTTS.AudioMarathon
🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs
Abstract
AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars:
long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.tadabur
Tadabur: A Large-Scale Quran Audio Dataset
The most comprehensive and richly annotated Qur'anic recitation corpus to date
Faisal Alherran
✦ Overview
Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur.mosel
Dataset Description, Collection, and Source
The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses.
In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.Bagpiper_PreTrain_Data
Bagpiper Pretraining Data
Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated
with Bagpiper, an open-ended audio language
model that learns bidirectional mappings between audio and comprehensive text
descriptions across speech, music, environmental sound, and mixtures.
The en metadata describes the primary rich-caption language. Source audio can
contain speech or singing in other languages; it is not an English-only audio
guarantee.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.StreamAudio-2M
StreamAudio-2M
Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a
stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips
are organised into six task subsets.
Subsets
Subset
Rows
Description
Stream_Audio_Understanding
90,738
Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA
Real_time_ASR
28,109
Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.ytseg
YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation
We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.UniST
UniST
This dataset contains UniST codec-token training data exported from local metadata and codec results.
We train UniSS with UniST data.
Schema
id: sample identifier
transcription: source transcription from metadata text
translation: qwen_trans, falling back to trans_text
source_glm, target_glm: GLM token lists
source_bicodec, target_bicodec: bicodec semantic token lists
bicodec_global: source bicodec global token list
dataset_name, src_lang, tgt_lang, split:… See the full description on the dataset page: https://huggingface.co/datasets/cmots/UniST.ParsVoice
ParsVoice
A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
📣 Accepted to the EMNLP 2026 Main Conference.
ParsVoice is the largest publicly available Persian speech–text corpus tailored for
training multi-speaker text-to-speech (TTS) systems. It is built from long-form
Persian audiobook recordings using a fully automated pipeline combining sentence-aware
segmentation, ASR transcription, a ParsBERT sentence-completion classifier, binary-search… See the full description on the dataset page: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice.2M-Belebele
2M-Belebele
Highly-Multilingual Speech and American Sign Language Comprehension Dataset
We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL).
The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.DeepDialogue-orpheus
DeepDialogue-orpheus
DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text.
🚨 Important Notice
This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.2M-Flores-ASL
2M-Flores
As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest
sentences in the original flores200 dataset.
To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded.
The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time.
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.danish-asr-leaderboard
Open Danish ASR Leaderboard — Results
Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models.
Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.risale-i-nur-sohbet
Risale-i Nur Sohbet
Prof. Dr. Şener Dilek’ten izin alındı.
Türkçe
Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde
birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış
sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri
kümelerine karıştırılmaz.
Kapsam
2095 sohbet, 954.66 saat 16 kHz mono FLAC ses
Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste
seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.everyayah-wav
everyayah-wav — Quranic recitation audio mirror
Full-mushaf Quranic recitation audio at 16 kHz mono 16-bit WAV, re-encoded
from everyayah.com for ML / ASR research.
This dataset is intentionally audio-only — no transcription text and no
alignment timings. The canonical Quranic text is widely available from
Tanzil and other public sources; pair this audio with
whatever text edition fits your use case.
Schema
Column
Type
Notes
audio
Audio(16000)
16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/dev-ahmedhany/everyayah-wav.hashed_data
Munch Hashed Index - Lightweight Audio Reference Dataset
📖 Overview
Munch Hashed Index is a lightweight reference dataset that provides SHA-256 hashes for all audio files in the Munch Urdu TTS Dataset. Instead of storing 1.27 TB of raw audio, this index stores only metadata and cryptographic hashes, enabling:
✅ Fast duplicate detection across 4.17 million audio samples
✅ Efficient dataset exploration without downloading terabytes
✅ Quick metadata queries (voice… See the full description on the dataset page: https://huggingface.co/datasets/humair025/hashed_data.kupe-asr-en-data
kupe-asr-en-mini-150m — data
Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly):
raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this.
mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this.
Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state.
from datasets import load_dataset
ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train")
parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.QuranTTS
QuranTTS v4
An ear-verified Quranic recitation corpus for TTS and speech restoration.
v4 supersedes the earlier v1_chunks / v2_clean / v2_raw configurations.
It is a full re-cut of the corpus at 24-bit / 48 kHz, with ayah boundaries
taken from forced alignment rather than fixed padding — which fixes the
truncated ghunnah endings present in earlier releases.
Clips
64,721
Duration
301.4 h
Reciters
16
Riwaya / style
Hafs, murattal
Coverage
112 surahs, 6… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/QuranTTS.Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.seamless-interaction-jefferson-annotations
Seamless Interaction Jefferson-Style Annotations
An automatic, turn-oriented annotation layer for the
Meta Seamless Interaction Dataset.
It compares the dataset's traditional transcript with an ASR-derived
Jefferson-style condition and supplies speech-act, communicative-purpose,
interactional-signal, alignment, and quality fields.
This is a derived noncommercial research dataset. It does not redistribute
the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.TIE_shorts
Dataset Card for TIE_Shorts
Dataset Summary
TIE_shorts is a derived version of the Technical Indian English (TIE) dataset, a large-scale speech dataset (~ 8K hours) originally consisting of approximately 750 GB of content
sourced from the NPTEL platform. The original TIE dataset contains around 9.8K technical lectures in English delivered by instructors from various regions across India,
with each lecture averaging about 50 minutes. These lectures cover a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/raianand/TIE_shorts.
