datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Granary
Granary: Speech Recognition and Translation Dataset in 25 European Languages
Granary is a large-scale, open-source multilingual speech dataset covering 25 European languages for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) tasks.
Overview
Granary addresses the scarcity of high-quality speech data for low-resource languages by consolidating multiple datasets under a unified framework:
🗣️ ~1M hours of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Granary.Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.StreamAudio-2M
StreamAudio-2M
Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a
stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips
are organised into six task subsets.
Subsets
Subset
Rows
Description
Stream_Audio_Understanding
90,738
Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA
Real_time_ASR
28,109
Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.quran-tajweed-phonetics
The complete phonetic layer of the Quran in the riwaya of Hafs 'an
'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every
phone carrying its tajweed attribution: madd class with its transmitted
length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt,
the seventeen sifat, and the rule that produced it.
Built and maintained by Quran Lab, a waqf building open technology in
the service of the Quran.
How it was built and verified
Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.HearInContextEnglish | 中文
HearInContext
A Benchmark for Implicit Context in Speech Recognition
Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative.
Same audio. Different contexts. Different meanings.
HearInContext is a Mandarin–English contextual speech recognition benchmark. It pairs the same audio with dialogue histories supporting different meanings to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/HearInContext.darija-asr-corpus
Darija ASR Corpus (dataset-core)
Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
This repo contains four source subsets: DODa, DVoice, Wiki, and
YouTube. Each subset carries its own upstream license/terms -- see below --
because they are drawn from four different original projects.
Subsets
Config
Rows
Audio bundled?
Upstream license
Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.chinese-lips-speech-slide-probe
Chinese-LiPS Speech + Slide Probe
A self-contained probe set for testing whether visual slide context helps
simultaneous speech translation — with the input as audio, not transcripts.
Why audio matters: feeding a transcript to a text LLM deletes the acoustic
ambiguity (homophones, polysemy) that slide context is meant to resolve; the
transcript already commits to one reading. Any honest test of "does vision help
streaming ST" must consume speech.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.bengali-talkshow-audio
Bengali Talkshow Audio Dataset
A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs.
Dataset Description
This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.pavo-bench
PAVO-Bench: 50K-Turn Benchmark for ASR-LLM-TTS Pipeline Routing
Code: github.com/vnmoorthy/pavo-bench · Paper: TMLR 2026 (accepted) · Authors: NarasingaMoorthy VeiluKanthaPerumal (UPenn), Mohammed Imthathullah (Google)
pip install git+https://github.com/vnmoorthy/pavo-bench.git
Headline results (vs fixed-cloud baseline, 50,000 voice turns)
Metric
Result
Significance
P95 end-to-end latency (H100, LibriSpeech)
−10.3% (−167 ms)
—
Median latency
−34%… See the full description on the dataset page: https://huggingface.co/datasets/vnmoorthy/pavo-bench.massive-yt-edu-queue
Massive YouTube Educational Video Queue
Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours.
Description
This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.MM-ContextASR-Bench
MM-ContextASR Bench
Metadata and evaluation splits for Multimodal Conversational Context for
LLM-Based ASR: Data Construction, Training, and Benchmark.
Dataset summary
Config
Examples
Audio
Context
Primary metric
mm_contextasr
1,250 (250 current utterances × 5 histories)
1,439 WAV files included
Controlled user-assistant dialogue
entity Recall
kespeech
19,212
Source ID only
Same-speaker speech and transcript
CER, SER, entity Recall
cv_yue
3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.open-vi-dialog-synthetic-100h
OpenDialog Vietnamese Synthetic Dialogue 100h
Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments.
12,000 chunks
30 seconds per chunk
100.0 hours total
Each item contains S1/S2 speaker labels, turn timings, target text,
relationship, pronouns, environment, topic, mood, and source reference IDs.
Audio renderer: vLLM-Omni VoxCPM2
Audio format: mono WAV, 48 kHz, 30 seconds per chunk
This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.multilingual-speech
Multilingual Indian Conversational Speech
A dataset of naturalistic, spontaneous two-speaker conversations across
13 Indian languages, with segment-level transcripts, speaker profiles,
timestamps, and recording metadata. Designed for ASR, TTS, speaker
diarization, and conversational speech research.
Languages (13)
Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali,
Odia, Punjabi, Tamil, Telugu, Urdu.
Content
Conversations… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/multilingual-speech.AASR-Bench
AASRBench
AASRBench is a bilingual audio benchmark for evaluating ASR output on spoken-language cleanup, correction, formatting, filtering, and rephrasing.
Contents
917 benchmark samples.
917 WAV files referenced by the benchmark records.
6,637 rubric questions across content, format, filter, and rephrase.
11 scenes: academic, customer_service, daily_chat, dictation_memo, explanation, meeting, navigation, passthrough, tech, vibe_coding, and voice_search.
510… See the full description on the dataset page: https://huggingface.co/datasets/Andrew0425/AASR-Bench.Aren
ARen — Arabic/English ASR Robustness Set
Curated and published by TNSA AI.
A small, deliberately hard evaluation set for Arabic and English speech
recognition. Every clip exists in three acoustic conditions so you can measure
not just how a model scores, but how fast it falls apart as the channel
degrades.
Built because clean read-speech benchmarks stop discriminating between modern
ASR systems long before real deployments stop breaking.
Why it exists
On clean… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/Aren.asr-benchmark-outputs
SaarAI ASR Benchmark Outputs
Raw per-utterance model outputs (transcription manifests) produced by the gsma-asr-bench runners on SaarAI/asr-leaderboard-datasets.
files: 501
utterances: 4374087
languages: 7
models: 47
Layout
data/<language_name>/<split>__<dataset_config>__<model_slug>.jsonl
index.jsonl # one record per file (language, split, model, rows, sha256, ...)
index.csv
Directories categorise by language name; the file name begins with the split name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-benchmark-outputs.hebrew-asr-vn
Hebrew ASR three-source training dataset
Derived from ivrit.ai crowd-transcribe-v5, crowd-recital and knesset-committees.
Original data and transcripts are credited to ivrit.ai and its contributors.
Pinned revisions and preparation rules are in metadata/sources.json and
metadata/preparation-config.json. VoxKnesset is excluded by user decision.
Source/split
Clips
Hours
crowd-recital/test
1,557
1.071
crowd-recital/train
45,372
33.258
crowd-recital/validation
1,033… See the full description on the dataset page: https://huggingface.co/datasets/benderrodriguez/hebrew-asr-vn.chinese-lips-longform-debug
Chinese-LiPS Long-Form (zh long streaming speech)
Reconstructed continuous long-speech streams from
BAAI/Chinese-LiPS, for
slide-aware / streaming speech-translation development and evaluation. Each
source video (one speaker, one scripted lecture with slides) was released as
pre-segmented clips; here they are re-joined into the full talk.
Two variants of the same 3 talks (~97 min speech total):
config
how segments are placed
use
orig_timeline
at their original session… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-longform-debug.pseudolabel-malaya-speech-stt-train-whisper-large-v3uk-pods
uk-pods - speech datasets of Ukrainian podcasts.
Preparation
Clone the dataset repository and extract the content of clips.tar.gz archive.
git clone https://huggingface.co/datasets/taras-sereda/uk-pods
cd uk-pods && tar -zxvf clips.tar.gz
To use these manifests for training/inference with NeMo [1] modify audio_filepath to absolute locations of audio files extracted in previous step.
# data_root=<clonned_repo_dir> # /home/ubuntu/uk-pods
data_root=$(realpath .)
sed -i… See the full description on the dataset page: https://huggingface.co/datasets/taras-sereda/uk-pods.soundscape-bench
SoundScape-Bench
200 held-out multilingual soundscapes with exact, automatically-gradable answer keys for evaluating
"universal audio annotation" — the task of describing everything audible in a clip (speech, who/when/
what/which-language/how-it-is-said, sound effects, music, and vocal bursts) as one structured JSON list.
It is the benchmark for the LAION Universal Audio Annotation Pipeline (UAAP).
Why it exists
Every clip is built by gluing together pieces we… See the full description on the dataset page: https://huggingface.co/datasets/laion/soundscape-bench.
