datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icelandic_asr
Icelandic ASR Collection
This repository collects six Icelandic speech corpora in directly loadable
Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a
convenience repackaging: the linked CLARIN-IS records and original dataset
repositories remain the canonical sources and should be cited when using the
data.
No configuration is selected by default. Choose a corpus configuration and,
for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.ToneWebinars
ToneWebinars
malayalam-asr-5K
Malayalam ASR 5K — Verified Anchor Set
5,225 manually verified Malayalam speech-transcript pairs (7.1 hours),
speaker-disjoint across train/dev/test. Every record in this release
carries is_verified: true — each transcript was checked, not machine-
generated and left unreviewed.
Splits (speaker-disjoint, source-aware)
split
utterances
%
disjoint units
speakers
hours
train
3,657
70%
9
7
4.8
dev
783
15%
5
5
1.1
test
785
15%
6
5
1.1
No speaker or… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-asr-5K.ToneSlavic
ToneSlavic
ToneSpeak
ToneSpeak
ToneBooks
ToneBooks
DialectalSpeech-ICL
DialectalSpeech-ICL
A speech-recognition dataset of African American English (AAE) utterances spanning
multiple regional varieties. Each record provides an audio clip, its verbatim
reference transcript, and speaker/region metadata intended for evaluating ASR and
in-context-learning approaches on dialectal, low-resource speech.
This release is a stratified sample of utterances drawn across all regional
collections.
Dataset Structure
Split
Utterances
test… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/DialectalSpeech-ICL.icu_medications
