datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MonsoonASR-Open-ASR-leaderboard-en-IN
Voice Arena Monsoon en-IN (public test)
Part of the Open ASR Leaderboard, in the main board's default column set, so it contributes to the headline Average WER for every model listed.
A conversational Indian English ASR test set that records who was speaking, not only what
was said. Every clip carries twelve speaker attributes — gender, age, native district
and state, education, occupation, income band, handset — so a difference between two
systems can be traced to a group of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN.MonsoonASR-Open-ASR-leaderboard-hi-IN
Voice Arena Monsoon hi (public test)
Part of the Open ASR Leaderboard, on the Multilingual tab, where a model is ranked only if it supports every selected language.
A conversational Hindi ASR test set that records who was speaking, not only what
was said. Every clip carries twelve speaker attributes — gender, age, native district
and state, education, occupation, income band, handset — so a difference between two
systems can be traced to a group of speakers instead of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-hi-IN.mongolian-stt-dataset
Mongolian Speech Dataset (v24 corpus)
Mongolian (Cyrillic Khalkha) read speech for ASR fine-tuning: 146.9 hours
across Common Voice v24, FLEURS, and MBSpeech.
2026-07-30 — two changes, read this if you pulled before that date.
YouTube-sourced audio removed. 598 clips (559 train / 39 validation, ~1.1 h)
are gone. Every remaining row is read speech from a redistributable public corpus.
This repo now hosts the v24 corpus. It previously held the v20 blend
(57,320 train / 3,017… See the full description on the dataset page: https://huggingface.co/datasets/Blgn94/mongolian-stt-dataset.mondegreen-asr-errors
Mondegreen ASR error pairs
(ASR hypothesis, gold text) pairs for Japanese ASR post-correction.
This build is simulated -- errors come from a phonetic corruption model, not from a real ASR system. It exists so the whole pipeline (gate training, benchmarks, figures, CI) is reproducible without a GPU. Treat every number derived from it as a stated assumption, not a measurement.
How it was made
synthetic text
-> phonetic corruption model (mondegreen.simulate)
->… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/mondegreen-asr-errors.vocal-money-codeswitch-asr-benchmark
Vocal Money — Yoruba–English Code-Switched ASR Benchmark
A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally
code-switched Yoruba–English speech, together with the reference transcriptions and the output of
every system on every clip, so that the published results can be recomputed or contradicted.
Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026.
Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.librispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/MONOKOTIL/librispeech_asr.mon_language_asr_audio
RFA Mon Language Voices
This dataset contains 14.8 hours of audio in the Mon language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Mon language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented into 3,634 manageable chunks and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mon_language_asr_audio.mon-voice-dataset
Dataset Summary
This dataset is a community-driven collection of the Mon language (ISO 639-3: mon). It contains 18,000+ sentences and corresponding voice recordings collected via a Mon keyboard application. The goal is to provide high-quality open-source data to support Mon language integration into global AI systems like Google Translate, OpenAI Whisper, and ChatGPT.
Supported Tasks
• Translation: Mon to English/Burmese/Pali.
• ASR (Speech-to-Text): For Mon voice… See the full description on the dataset page: https://huggingface.co/datasets/Nenemin95/mon-voice-dataset.Czech-Speech-Monospeaker-Honza
Important
This dataset comes from voxpopuli.
We selected the most frequent male speaker in the dataset and created a separate single-speaker dataset.
Processing performed:
Recording of a neutral speaker, in large quantities
Denoising with https://huggingface.co/speechbrain/sepformer-whamr16k
voiceenv-money-transfer
VoiceEnv: Money Transfer
An autonomously-extracted voice-agent RL environment built from a single real
human-human call recording (HVB dataset). Generated by
VoiceEnv.
Contents
env.yaml — the full environment spec (task, persona, tools, rubric, expert reference)
expert_reference/source_call.wav — the original real human-human call (anchor for grounded judging)
caller_clips/ — per-turn caller audio slices (real human voice, for stateless eval)
human_response_clips/ — what… See the full description on the dataset page: https://huggingface.co/datasets/karthik/voiceenv-money-transfer.mondegreenbench
MondegreensEval
Audio companion dataset for MondegreensEval: A Phonetic Benchmark for Measuring
Language-Model Bias in Automatic Speech Recognition (ICML ML for Audio Workshop, 2026).
Code, evaluation pipeline, and per-model transcription/metric outputs:
https://github.com/soarhigh/mondegreenbench
Mondegreens — phonetically near-identical phrase pairs with distinct meanings — expose a
measurable failure mode in decoder-based ASR: the model's internal language-model prior
can… See the full description on the dataset page: https://huggingface.co/datasets/soarhigh/mondegreenbench.17-minute-world-languages_mongol
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/mongole/
Site à scrapper
mon-words-dataset
Dataset Summary
This dataset is a community-driven collection of the Mon language (ISO 639-3: mon). It contains 15,000+ sentences and corresponding voice recordings collected via a Mon keyboard application. The goal is to provide high-quality open-source data to support Mon language integration into global AI systems like Google Translate, OpenAI Whisper, and ChatGPT.
Supported Tasks
• Translation: Mon to English/Burmese/pali.
• ASR (Speech-to-Text): For Mon voice… See the full description on the dataset page: https://huggingface.co/datasets/Nenemin95/mon-words-dataset.
