CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ccmusic-database /CTIS Dataset Card for Chinese Traditional Instrument Sound Original Content The original dataset is created by [1], with no evaluation provided. The original CTIS dataset contains recordings from 287 varieties of Chinese traditional instruments, reformed Chinese musical instruments, and instruments from ethnic minority groups. Notably, some of these instruments are rarely encountered by the majority of the Chinese populace. The dataset was later utilized by [2] for Chinese… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/CTIS.audioaudio-classification10K<n<100K29 likes849 downloads7mo agoHugging Face02ctaguchi /d2l4asr-wiki-jaaudio100K<n<1M0 likes160 downloads22d agoHugging Face03ctaguchi /d2l4asr-wiki-en_audioaudio100K<n<1M0 likes118 downloads28d agoHugging Face04ctaguchi /ikema_youtube_asr_full_with_longaudio10K<n<100K0 likes108 downloads3mo agoHugging Face05ctaguchi /SLR35_javaneseaudio100K<n<1M0 likes106 downloads23d agoHugging Face06ctaguchi /killkan Killkan: Speech Recognition dataset for Kichwa Killkan (Kichwa uyachkata payllatak killkak anta) is the first automatic speech recognition (ASR) dataset for the Kichwa language. See also our paper (https://arxiv.org/abs/2404.15501). audioautomatic-speech-recognition1K<n<10K0 likes99 downloads2y agoHugging Face07touringwithayo /nollywood-ctc-scored-ep3-hauwaaudion<1K1 likes93 downloads8d agoHugging Face08ctaguchi /d2l4asr-wiki-enaudio10K<n<100K0 likes86 downloads1mo agoHugging Face09DevVault /ctaudio1K<n<10K0 likes68 downloads10d agoHugging Face10bobboyms /phoneme-ctc-spanish-52h-noisyaudio10K<n<100K0 likes54 downloads9mo agoHugging Face11bobboyms /phoneme-ctc-english-60haudio10K<n<100K0 likes46 downloads9mo agoHugging Face12bobboyms /phoneme-ctc-english-41haudio10K<n<100K0 likes41 downloads11mo agoHugging Face13akmalsultanov /usc_cleaned_ctc_filteredgatedaudio10K<n<100K0 likes38 downloads1y agoHugging Face14ctaguchi /ikema_dict_asraudio1K<n<10K0 likes31 downloads11mo agoHugging Face15DynamicSuperb /SingingVoiceDeepfakeDetection_CtrSVDD_ACEKiSing_M4Singeraudion<1K1 likes27 downloads2y agoHugging Face16bobboyms /phoneme-ctc-english-60h-balanced Phoneme CTC — English 60h (Balanced & Normalized) A cleaned, normalized and phoneme-balanced version of bobboyms/phoneme-ctc-english-60h-noisy, for training phoneme recognition models (CTC) — e.g. as the native acoustic model behind pronunciation-feedback systems. What's different from the source dataset Label noise removed Roman numerals dropped — eSpeak reads ii/iv/… as "Roman two/four", producing labels that don't match the audio. Non-English phonemes dropped… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/phoneme-ctc-english-60h-balanced.audioautomatic-speech-recognition10K<n<100K0 likes27 downloads3mo agoHugging Face17ctam8736 /papi_asr_testaudio10K<n<100K0 likes23 downloads3y agoHugging Face18ctam8736 /papi_asraudio10K<n<100K0 likes21 downloads3y agoHugging Face19ctaguchi /ikema_dictionary_examples_datasetaudio1K<n<10K0 likes18 downloads6mo agoHugging Face20ctam8736 /papi_asr_miniaudion<1K0 likes15 downloads3y agoHugging Face21bobboyms /phoneme-ctc-english-60h-noisyaudio10K<n<100K0 likes13 downloads9mo agoHugging Face22apple121 /MMAU-Pro-Ctrlaudio1K<n<10K0 likes13 downloads4mo agoHugging Face23ctaguchi /ikema_youtube_asr_testaudion<1K0 likes12 downloads11mo agoHugging Face24AhunInteligence /hubert_ctc_ftaudio1K<n<10K0 likes8 downloads8mo agoHugging Face25NancyT /wur_ctc_kln_scoredaudion<1K0 likes7 downloads3mo agoHugging Face26heimayuan /kws_dataset_ctgated ygyuan/kws_dataset_ct Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: train: 941 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/heimayuan/kws_dataset_ct.audioaudio-classification0 likes6 downloads2mo agoHugging Face27IAmNotAnanth /sinhala-ctc-111hgated Sinhala ASR – Consolidated OpenSLR (SLR52) Dataset Summary This dataset is a consolidated and cleaned version of the Sinhala Automatic Speech Recognition (ASR) dataset from OpenSLR (SLR52). The original OpenSLR release distributes the data across multiple subsets (0–9, a–f). This repository merges all subsets into a single unified dataset containing approximately 111 hours of speech audio. Dataset Description Consolidation All OpenSLR SLR52… See the full description on the dataset page: https://huggingface.co/datasets/IAmNotAnanth/sinhala-ctc-111h.audioautomatic-speech-recognition100K<n<1M0 likes5 downloads9mo agoHugging Face28ygyuan /kws_testset_ct_sphgated ygyuan/kws_testset_ct_sph Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 1 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_ct_sph.audioaudio-classification0 likes5 downloads2mo agoHugging Face29ctem049 /diff_dataaudion<1K0 likes4 downloads3y agoHugging Face30om-ai /vn-speech-text-ctv-v1.1gatedaudio10K<n<100K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.