CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aalto-Speech-Synthesis /icelandic_asr Icelandic ASR Collection This repository collects six Icelandic speech corpora in directly loadable Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a convenience repackaging: the linked CLARIN-IS records and original dataset repositories remain the canonical sources and should be cited when using the data. No configuration is selected by default. Choose a corpus configuration and, for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.audioautomatic-speech-recognition1M<n<10M0 likes592 downloads21d agoHugging Face02ICML-2026 /ToneWebinars ToneWebinars audioautomatic-speech-recognition100K<n<1M0 likes276 downloads8mo agoHugging Face03icfoss /malayalam-asr-5K Malayalam ASR 5K — Verified Anchor Set 5,225 manually verified Malayalam speech-transcript pairs (7.1 hours), speaker-disjoint across train/dev/test. Every record in this release carries is_verified: true — each transcript was checked, not machine- generated and left unreviewed. Splits (speaker-disjoint, source-aware) split utterances % disjoint units speakers hours train 3,657 70% 9 7 4.8 dev 783 15% 5 5 1.1 test 785 15% 6 5 1.1 No speaker or… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-asr-5K.audioautomatic-speech-recognition1K<n<10K0 likes94 downloads10d agoHugging Face04ICML-2026 /ToneSlavic ToneSlavic audioautomatic-speech-recognition1M<n<10M0 likes56 downloads8mo agoHugging Face05ICML-2026 /ToneSpeak ToneSpeak audioautomatic-speech-recognition1K<n<10K0 likes32 downloads8mo agoHugging Face06ICML-2026 /ToneBooks ToneBooks audiotext-to-speech10K<n<100K1 likes18 downloads8mo agoHugging Face07CentificAIResearch /DialectalSpeech-ICL DialectalSpeech-ICL A speech-recognition dataset of African American English (AAE) utterances spanning multiple regional varieties. Each record provides an audio clip, its verbatim reference transcript, and speaker/region metadata intended for evaluating ASR and in-context-learning approaches on dialectal, low-resource speech. This release is a stratified sample of utterances drawn across all regional collections. Dataset Structure Split Utterances test… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/DialectalSpeech-ICL.audioautomatic-speech-recognitionn<1K2 likes15 downloads3mo agoHugging Face08arnepeine /icu_medicationstextautomatic-speech-recognitionn<1K0 likes9 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.