datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-asr-leaderboard-resultsasr-jargon-specialized-vocabulary
A Dataset for Evaluating ASR on Specialized Vocabulary
Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026).
Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code
Configs
Config
Language
Description
synthetic_terms_en
English
Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms
synthetic_terms_pt
Portuguese
Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.asr_benchmark_storeclipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.ASRD-Dataset
Adversarial Surface-Form Robustness Dataset (ASRD)
Anonymous Repository for Double-Blind ReviewNeurIPS 2026 Workshop
1. Dataset Overview
Standard safety evaluations of large language models routinely measure model refusal and compliance using canonical plain-text instructions. However, deployed systems frequently encounter non-canonical inputs containing expressive symbols (emojis), character-level substitutions (homoglyphs, leetspeak), structured encodings… See the full description on the dataset page: https://huggingface.co/datasets/asrd-research/ASRD-Dataset.uyghur-ASR-dataset
Uyghur ASR Corpus (Latin Transliteration)
A speech corpus for Uyghur automatic speech recognition, with transcriptions in a
case-sensitive Latin transliteration scheme. Approximately 23 hours of audio across
9,468 clips.
Uyghur is a Turkic language spoken by roughly 10–12 million people. It is severely
under-represented in open speech datasets, and this corpus is intended to support ASR research
for the language.
Dataset summary
Language
Uyghur (ug)… See the full description on the dataset page: https://huggingface.co/datasets/Shramadeepd/uyghur-ASR-dataset.polish-tedx-asr-eval
Polish-TEDx-ASR-Eval
A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks.
Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.Kinyarwanda_Engligh_Multilingual_ASRThis dataset was created from Mozilla's Common Voice dataset for the purposes of Multilingual ASR on Kinyarwanda and English.
The dataset contains 3000 hours of multilingual training samples, 300 hours of validation samples and 200 of testing samples.
somali-asr-synthetic-youtube
Somali ASR Synthetic YouTube Dataset
A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping.
Dataset Summary
Split
Samples
train
~4,393
validation
200
test
100
Total
~4,693
Language: Somali (so)
Audio format: WAV, 16 kHz, mono, 16-bit PCM
Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.almaz-asr-roster
The ALMAZ ASR Roster
Archived at Zenodo: 10.5281/zenodo.22761871
(concept DOI, always resolves to the latest version).
A curated catalog of Azerbaijani speech-to-text artifacts: corpora, models,
services, benchmarks and tools. Companion to the
ALMAZ Resource Roster,
which does the same for text.
Schema matches the text roster so the two join, plus three columns speech needs
and text does not: hours, condition, and verified.
The verified column
The standard way a… See the full description on the dataset page: https://huggingface.co/datasets/almaz-nlp/almaz-asr-roster.Swedia-ASR-Dataset
Swedia ASR Dataset
This repository contains a small Swedish ASR evaluation dataset based on speech
transcriptions from Swedia 2000. It was assembled to compare automatic
speech-recognition output against manually corrected reference transcriptions
for Swedish dialectal speech.
The dataset is useful for quick experiments with Swedish ASR systems, especially
when you want to inspect recognition quality on spontaneous speech from
different regions, speakers, ages, and genders.… See the full description on the dataset page: https://huggingface.co/datasets/kvest/Swedia-ASR-Dataset.CORAA-MUPE-ASRsovereign-asr-bench
Sovereign ASR Bench — RTX 5090
Local, self-hosted automatic speech recognition benchmarks on one RTX 5090 32GB.
Part of the WITCHEER local-AI rig. Methodology that matters: load-once measurement (so
RTFx times transcription, not model load), one shared text normalizer applied to every
model output and reference, and micro-averaged WER (total errors / total reference
words — the LibriSpeech standard). The board lives as data in board.csv (shown in the viewer).
Board —… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/sovereign-asr-bench.ben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.2026-dwesui-g01-neurologia
DWESUI 2026 - Grupa 1 - neurologia
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 1 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Domena: neurologia
Licencja zrodla: nagrania YouTube CC-BY + synteza TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.ayurvedic-remedies
Ayurvedic Remedies Dataset
This dataset contains curated mappings of over 100 common symptoms and diseases to Ayurvedic remedies. It is designed to support fine-tuning language models or powering AI chatbots with traditional Indian medical knowledge.
Dataset Structure
Column
Description
Symptom
The name of the symptom, condition, or disease
Remedy
One or more Ayurvedic treatments or herbal/home remedies
Each row provides a traditional Ayurvedic solution… See the full description on the dataset page: https://huggingface.co/datasets/ASR01/ayurvedic-remedies.2026-dwesui-g02-kulinarna
DWESUI 2026 - Grupa 2 - kulinarna (PIEROGA)
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 2 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna
Domena: kulinarna
Licencja zrodla: nagrania YouTube CC-BY/CC-BY-SA + TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g02-kulinarna.asr_bangla_2024open-asr-leaderboard-evals-alluzbek-legal-datasetresponses-and-asr-labels-small-models
LLM Responses and ASR Labels — Small Models
Model responses to harmful prompts, labelled by 4 LLM-as-judge guards.Companion dataset for the master's thesis ASR Signal Geometry: Dense Representations vs. SAE Features (HSE, 2025).
Dataset composition
N = 4 326 prompts per model, (no adversarial suffix). Two sources:
Source
N
Description
JailbreakBench ()
100
Curated harmful behaviours
Anthropic HH-RLHF red-team-attempts ()
4 226
Red-team conversations… See the full description on the dataset page: https://huggingface.co/datasets/SabrinaSadiekh/responses-and-asr-labels-small-models.ASRS-ChatGPT
Dataset Summary
The dataset contains a total of 9984 incident records and 9 columns. Some of the columns contain ground truth values whereas others contain information generated by ChatGPT based on the incident Narratives.
The creation of this dataset is aimed at providing researchers with columns generated by using ChatGPT API which is not freely available.
Dataset Structure
The column names present in the dataset and their descriptions are provided below:
Column… See the full description on the dataset page: https://huggingface.co/datasets/archanatikayatray/ASRS-ChatGPT.dataset_mistral_instruct_commerce_operations-performance_best_practices_v1Norm_Malayalam_Evaluation_samples
Malayalam ASR Reference Prediction dataset
This repository contains evaluation results from the Malayalam ASR model "vrclc/Whisper_small_malayalam" using the "google/fleurs" dataset.
ASR Model Name: vrclc/Whisper_small_malayalam
Dataset: google/fleurs
Curated by: VRCLC
vrclc/Whisper_small_malayalam was trained with 50 hours of Malayalam speech data.
The test set of google/fleurs dataset which consists of Malayalam speech data was used to evaluate the model
The evaluation of 500… See the full description on the dataset page: https://huggingface.co/datasets/asr-malayalam/Norm_Malayalam_Evaluation_samples.Devnagari-ASRASRLive.csvFlores-subsetEDEN_ASR_Data
EDEN ASR Dataset
A subset of this data was used to support the development of empathetic feedback modules in EDEN and its prior work.
The dataset contains audio clips of native Mandarin speakers. The speakers conversed with a chatbot hosted on an English practice platform.
3081 audio clips from 613 conversations and 163 users remained after filtering.
The filtering process removes audio clips containing only Mandarin, duplicates, and a subset of self-introductions from the users.… See the full description on the dataset page: https://huggingface.co/datasets/sylviali/EDEN_ASR_Data.open-asr-leaderboard-evals-ex-cvAmharic_ASR_Dataset
