datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
asr-jargon-specialized-vocabulary
A Dataset for Evaluating ASR on Specialized Vocabulary
Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026).
Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code
Configs
Config
Language
Description
synthetic_terms_en
English
Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms
synthetic_terms_pt
Portuguese
Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.clipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.uyghur-ASR-dataset
Uyghur ASR Corpus (Latin Transliteration)
A speech corpus for Uyghur automatic speech recognition, with transcriptions in a
case-sensitive Latin transliteration scheme. Approximately 23 hours of audio across
9,468 clips.
Uyghur is a Turkic language spoken by roughly 10–12 million people. It is severely
under-represented in open speech datasets, and this corpus is intended to support ASR research
for the language.
Dataset summary
Language
Uyghur (ug)… See the full description on the dataset page: https://huggingface.co/datasets/Shramadeepd/uyghur-ASR-dataset.polish-tedx-asr-eval
Polish-TEDx-ASR-Eval
A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks.
Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.Swedia-ASR-Dataset
Swedia ASR Dataset
This repository contains a small Swedish ASR evaluation dataset based on speech
transcriptions from Swedia 2000. It was assembled to compare automatic
speech-recognition output against manually corrected reference transcriptions
for Swedish dialectal speech.
The dataset is useful for quick experiments with Swedish ASR systems, especially
when you want to inspect recognition quality on spontaneous speech from
different regions, speakers, ages, and genders.… See the full description on the dataset page: https://huggingface.co/datasets/kvest/Swedia-ASR-Dataset.somali-asr-synthetic-youtube
Somali ASR Synthetic YouTube Dataset
A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping.
Dataset Summary
Split
Samples
train
~4,393
validation
200
test
100
Total
~4,693
Language: Somali (so)
Audio format: WAV, 16 kHz, mono, 16-bit PCM
Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.ben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.2026-dwesui-g01-neurologia
DWESUI 2026 - Grupa 1 - neurologia
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 1 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Domena: neurologia
Licencja zrodla: nagrania YouTube CC-BY + synteza TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.2026-dwesui-g02-kulinarna
DWESUI 2026 - Grupa 2 - kulinarna (PIEROGA)
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 2 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna
Domena: kulinarna
Licencja zrodla: nagrania YouTube CC-BY/CC-BY-SA + TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g02-kulinarna.chuvash-asr-final
Chuvash ASR1
Небольшой аудиодатасет для распознавания чувашской речи.
Structure
audio/ — аудиофайлы
train.csv — метаданные с колонками:
file_name
text_chv
client_id
Columns
file_name — относительный путь к аудиофайлу
text_chv — расшифровка на чувашском языке
client_id — идентификатор говорящего
Example
file_name,text_chv,client_id
audio/00001.ogg,Салам,spk01
audio/00002.ogg,Ырӑ ир,spk01
Web demo (single page)
Файл web_app.py… See the full description on the dataset page: https://huggingface.co/datasets/yulia774/chuvash-asr-final.
