datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voicehub-arena-seed-tts-eval
VoiceHub Arena — full English Seed-TTS-Eval
35,904 synthesized WAV files: 33 model families × the same 1,088 target texts.
The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB.
All 198 shards and every WAV SHA256 were verified after backup.
Interactive leaderboard and all audio samples
· Source repository (access required).
Contents
audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.emova-asr-tts-eval
EMOVA-ASR-TTS-Eval
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-ASR-TTS-Eval is a dataset designed for evaluating the ASR and TTS performance of Omni-modal LLMs. It is derived from the test-clean set of the LibriSpeech dataset. This dataset is part of the EMOVA-Datasets collection. We extract the speech units using the EMOVA Speech Tokenizer.
Structure
This… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-asr-tts-eval.nb-asr-eval-withwav-sorted
NB-ASR Eval with WAV Sorted
Hardest-first copy of NbAiLab/nb-asr-eval-withwav for targeted human cleanup.
Rows are intended to be sorted independently within each split by ASR/WER difficulty.
Audio paths and split metadata layout are preserved so downstream tools can switch
from the original repo to NbAiLab/nb-asr-eval-withwav-sorted without changing file lookup logic.
After scoring, each metadata row may include original_source_index, priority_rank,
asr_wer, asr_cer… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-eval-withwav-sorted.MLC-SLM-Eval
Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) Eval Groundtruth
🖥️ Overview
In the MLC-SLM challenge, we only provided the participants with the audio files of the Eval sets.
Now, we release the oracle segmentation, speaker labels, and transcriptions of the Eval sets to facilitate further research by all participants on the MLC-SLM dataset!
In addition, the MLC-SLM challenge summary paper "Summary on The Multilingual Conversational Speech… See the full description on the dataset page: https://huggingface.co/datasets/bsmu/MLC-SLM-Eval.open-ko-s2s-eval-artifacts
Open Ko-S2S 평가 산출물 (감사용)
⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다
KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의
ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과
채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다.
Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라
ref 원문을 그대로 담고 있습니다.
라이선스 보유자의 검증 절차
AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다.
리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다.
자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.vistaar_small_asr_eval
Vistaar Small ASR Eval
Dataset Description
The Vistaar Small ASR Eval dataset is a multilingual automatic speech recognition evaluation dataset containing 9,486 audio samples across 12 Indian languages. This dataset represents a subset of the larger Vistaar dataset published by AI4Bharat, designed specifically for evaluating ASR model performance on diverse Indian language speech data. A smaller evaluation dataset was created for the use-cases where a quick benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/vistaar_small_asr_eval.CoSHE-Eval
🎙️ CoSHE-Eval: A Code-Switching ASR Benchmark for Hindi–English Speech
🧠 Overview
CoSHE-Eval is an evaluation dataset curated for testing Automatic Speech Recognition (ASR) systems on Hindi-English code-mixed speech.It focuses on bilingual conversational contexts commonly found in India, where Hindi (in Devanagari) and English (in Latin script) co-occur naturally within the same utterance.
Detailed Blog: CoSHE-Eval Blog
Technical Specifications… See the full description on the dataset page: https://huggingface.co/datasets/soketlabs/CoSHE-Eval.agri-voice-eval
Agricultural Voice Evaluation Set — Hindi, Telugu, Odia
Human quality-checked FarmerChat field recordings with per-clip audio and human reference transcripts,
used to evaluate Digital Green's agricultural voice pipeline stage by stage. Companion to the larger
Agri STT Benchmarking Dataset,
focused on the multi-speaker / noisy conditions that motivate speaker selection and enhancement.
Each clip carries two references: the farmer's own (main-speaker) transcript — the… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/agri-voice-eval.eval-whatsapp
Dataset Card for ivrit.ai Whatsapp Eval
Evaluation dataset of Hebrew Whatsapp voice-messages, expert-transcribed.
Dataset Details
Dataset Description
This dataset containeד Whatsapp hebrew voice recordings gathered around April 2025.
The recordings are of volunteer native hebrew speakers using consumer devices in natural environments.
Each recording is by a single speaker about a random topic of their choice spoken in a non-scripted natural manner.
Each… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-whatsapp.LEMAS-Dataset-eval
Overview
This dataset is part of LEMAS-Project(lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-eval.ksponspeech-evalpaper link: https://www.mdpi.com/846876
Indic_ASR_Eval
Indic ASR Eval
A curated evaluation set for Indic-language automatic speech recognition.
100 samples are sampled (seed = 42) from each (source dataset × language)
cell of seven public Indic ASR corpora. Each source corpus is published
as its own dataset config with a single test split, at 16 kHz.
Rows: 6,169 across 7 configs
Total audio: ~13.3 hours
Sampling rate: 16 kHz (mono)
Split: test (single split in every config)
Configs
Config
Rows
Notes
kathbath… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/Indic_ASR_Eval.nvidia-brain-noise-evaluation-dataset
Nvidia Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 64 samples organized across multiple splits and 32 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 2 samples
test: 2 samples
noisy-bg-snr-20: 2 samples
test: 2 samples
noisy-bg-snr-30: 2 samples
test: 2 samples
noisy-bg-snr-40: 2 samples
test: 2 samples
noisy-bg-snr-50: 2 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/nvidia-brain-noise-evaluation-dataset.tiron-eval-meetings
Tiron evaluation meetings
The 17 held-out whole meetings used for the benchmarks on the
Trelis/tiron model card — far-field
single-channel audio (16 kHz mono WAV) with reference speaker-attributed
transcripts, packaged so results can be reproduced with the
Tiron harness.
split
meetings
source
ami
ES2004a, IS1009a, TS3003a, EN2002a
AMI Meeting Corpus, single distant microphone (Array1-01)
icsi
Bmr013, Bmr018, Bro021
ICSI Meeting Corpus, mean of 4 distant PZM room… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/tiron-eval-meetings.polish-tedx-asr-eval
Polish-TEDx-ASR-Eval
A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks.
Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.bg-med-consultations-eval
Bulgarian Medical Consultations — evaluation set
156 synthesised Bulgarian doctor–patient consultations with exact ground
truth: 133.4 minutes, 2,544 turns, 11 native Bulgarian voices.
The reference is not an annotation. Every turn was placed on the timeline by the
generator, so the RTTM is a construction — correct by definition, with none of
the annotator disagreement that inflates published DER.
What is in it
directory
contents
audio/
16 kHz mono… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bg-med-consultations-eval.Audio-Understanding-Bitrate-Eval-0426
Audio Understanding — MP3 Bitrate Evaluation (April 2026)
Empirical eval measuring how MP3 compression bitrate affects transcription accuracy across every audio-input LLM available on OpenRouter.
📝 Blog post: MP3 Bitrate Sensitivity in Audio-Multimodal LLMs
💻 Code & methodology: github.com/danielrosehill/Audio-Understanding-Bitrate-Eval-0426
TL;DR
Ran a benchmark across 12 OpenRouter audio-multimodal models × 4 dictation samples × 5 MP3 bitrates (16/24/32/48/64 kbps)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Bitrate-Eval-0426.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.Small-STT-Eval-Audio-Dataset
Small STT Eval Audio Dataset
A small speech-to-text evaluation dataset containing 92 audio samples with ground truth transcriptions. Designed for evaluating STT systems on technical vocabulary, code-switching (English/Hebrew), and various speaking styles.
Dataset Description
This dataset contains audio recordings with accompanying transcriptions across multiple categories:
Category
Count
Description
tech_github
5
GitHub-related technical vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Small-STT-Eval-Audio-Dataset.atc-asr-eval
ATC ASR Evaluation Set
Human-verified air-traffic-control transmissions for evaluating ASR models.
Each clip is 16 kHz mono WAV with a corrected ground-truth transcript in
metadata.csv (columns: file_name, transcription, icao, city, region, source, feed).
Audio captured from LiveATC.net feeds. Private — not for redistribution
(LiveATC terms prohibit rebroadcasting).
Albayzin-2024-BBS-S2T-eval
Albayzin 2024 Bilingual Basque-Spanish Speech to Text (BBS-S2T) Challenge - Evaluation dataset
see Albayzin_2024_BBS-S2T_EvalPlan for a description of the challenge.
This is the evaluation data for the challenge.
The database consists of a single split:
eval : 12498 audio segments
How to download this database
1 - If you can handle yourself comfortably with Huggingface Datasets:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/gttsehu/Albayzin-2024-BBS-S2T-eval.bocalantics-wolof-eval
Bocalantics Wolof eval set
The held-out Wolof test rows from MOH749/Bocalantics-2.0, with audio attached, so a
candidate model can be scored without re-materialising anything.
3,557 rows, 5.38 hours.
Why it exists
"Beats unadapted Whisper" is not a result for Wolof. Whisper has seen 2 of the parent
corpus's 26 languages, so beating it is arithmetic rather than evidence. The bar is the
best published model for the language -… See the full description on the dataset page: https://huggingface.co/datasets/MOH749/bocalantics-wolof-eval.2025-zwesui-g02-medyczna
ZWESUI 2025 - Grupa 2 - medyczna (mowa syntetyczna)
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2025 (pierwsza), tryb niestacjonarny.
Zespol (atrybucja): Grupa 2 (2025)
Zrodlo oryginalne: https://huggingface.co/datasets/yanvoi/med_male_female_r2
Domena: medyczna
Opis: 200 zdań z terminologią medyczną, mowa syntetyczna (ElevenLabs), głosy męskie i żeńskie; zbiór użyty… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2025-zwesui-g02-medyczna.Bam_ASR_Eval_500
Bam_ASR_Eval_500 Dataset
Dataset Description
Bam_ASR_Eval_500 is a curated evaluation dataset for Automatic Speech Recognition (ASR) models in Bambara (Bamanakan), a major language spoken in Mali and West Africa. This dataset comprises 500 audio recordings totaling approximately 36.69 minutes of annotated speech, designed specifically for benchmarking ASR systems. It focuses on real-world challenges in low-resource languages like Bambara, including spontaneous speech… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/Bam_ASR_Eval_500.Indic-subtitler-audio_evals
Indic_audio_evals
As part of this project. We are evaluating our performance of various ASR models as well
in a benchmarking dataset, we have created in various languages. This benchmarking dataset
is more alligned to real-world use-cases rather than having any academic datasets.
About Dataset
Dataset Link in HuggingFace: kurianbenoy/Indic-subtitler-audio_evals
This dataset contains audio file in .wav format and video file in .mp4. The respective groundtruth will be… See the full description on the dataset page: https://huggingface.co/datasets/kurianbenoy/Indic-subtitler-audio_evals.nyana-eval
Nyana-Eval Dataset
Dataset Description
Nyana-Eval is a compact, stratified evaluation subset for benchmarking Automatic Speech Recognition (ASR) models in Bambara. It consists of 45 audio recordings totaling approximately 3.03 minutes, carefully selected to represent real-world linguistic and acoustic challenges in low-resource Bambara speech. This dataset is derived from the larger RobotsMali/Bam_ASR_Eval_500 corpus and is optimized for quick, reproducible human… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/nyana-eval.aura-phone-dictation-eval
Aura Phone Dictation Eval
Evaluation set of 365 progressive audio clips from 142 phone-number dictation sequences extracted from Aura Hindi/English call-center recordings.
This dataset is used to evaluate end-of-turn (EOT) detection models on structured phone-number dictation. Each sequence captures a caller dictating a 10-digit Indian mobile number across multiple speech segments. Progressive clips accumulate earlier segments plus trailing silence, ending with a final clip once… See the full description on the dataset page: https://huggingface.co/datasets/ananth-r-gnani/aura-phone-dictation-eval.YouTube-Evaluation-Set
Awaaz se Alfaaz — YouTube Evaluation Set
This dataset is the realistic multi-speaker evaluation set used in Awaaz se Alfaaz, accepted at LaTeLL 2026 — "Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing." It contains 30 short-form Urdu YouTube videos (YouTube Shorts) covering a mix of news, sports, and current affairs content, along with human annotated gold transcripts and transcripts produced by… See the full description on the dataset page: https://huggingface.co/datasets/awaaz-se-alfaaz/YouTube-Evaluation-Set.asr-evaluations
