CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TigreGotico /FalaBracarense_splitsdataset website: projectofalabracarense Licence CC - BY - NC - ND Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives audioautomatic-speech-recognition100K<n<1M0 likes2.2k downloads1y agoHugging Face02kazeric /OpenBible_Swahili_book_splitaudio10K<n<100K0 likes1.2k downloads2y agoHugging Face03kawsarahmd /common_voice_13_0_bn_multi_splitaudio1M<n<10M0 likes634 downloads2y agoHugging Face04NickWeng /tw_parliament_split Parliament Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the zh_tw split of disco-eth/WorldSpeech. It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD parliamentary proceedings. This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style text filtering pipeline. Audio is preserved from the upstream dataset and cast as a Hugging Face Audio(sampling_rate=24000) feature. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tw_parliament_split.audioautomatic-speech-recognition100K<n<1M0 likes529 downloads8d agoHugging Face05hhim8826 /japanese-anime-speech-v2-split-150k japanese-anime-speech-v2-split-150k joujiboi/japanese-anime-speech-v2 的 150,000 筆隨機子集,已切好 train / test。 split rows train 135,000 test 15,000 total 150,000 Columns audio — 16 kHz mp3,與原始資料完全相同(未重新編碼) sentence — 轉錄文字(原始欄位名為 transcription) How it was built 來源的 sfw(271,788 筆)與 nsfw(20,849 筆)兩個 split 都有使用,並依原始比例分配名額 (sfw 139,313 / nsfw 10,687)。 在每個 split 的全部列上做無放回均勻抽樣,因此每筆資料被選中的機率相同。 抽出後整體打亂,前 15,000 筆為 test,其餘為 train,train / test… See the full description on the dataset page: https://huggingface.co/datasets/hhim8826/japanese-anime-speech-v2-split-150k.audioautomatic-speech-recognition100K<n<1M0 likes434 downloads2mo agoHugging Face06AdrienB134 /Emilia-dataset-french-splitaudio100K<n<1M4 likes432 downloads2y agoHugging Face07ghanaopenai /ghana-female-twi-speech-asr-8word-splits This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 8-Word Speech Segments 51139 speech-text pairs split from 30-min recordings. Processing pipeline Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-female-twi-speech-asr-8word-splits.audioautomatic-speech-recognition10K<n<100K0 likes424 downloads3mo agoHugging Face08sanchit-gandhi /earnings22_robust_splitfrom datasets import load_dataset, DatasetDict ds = load_dataset("anton-l/earnings22_robust", split="test") print(ds) print("\n", "Split to ==>", "\n") # split train 90%/ dev 5% / test 5% # split twice and combine train_devtest = ds.train_test_split(shuffle=True, seed=1, test_size=0.1) dev_test = train_devtest['test'].train_test_split(shuffle=True, seed=1, test_size=0.5) ds_train_dev_test = DatasetDict({'train': train_devtest['train'], 'validation': dev_test['train'], 'test':… See the full description on the dataset page: https://huggingface.co/datasets/sanchit-gandhi/earnings22_robust_split.audio10K<n<100K0 likes384 downloads4y agoHugging Face09rmndrnts /wikipedia_asr_splittedSynthetic dataset based on highly specialised texts from Wikipedia. This version is splitted by category. Voiced using Yandex SpeechKit with random voices, roles and speech rate. Can be used to evaluate ASR models not trained on given domains and to identify areas that the model does not handle well. Non-splitted version can be found here: https://huggingface.co/datasets/rmndrnts/wikipedia_asr audio100K<n<1M0 likes341 downloads2y agoHugging Face10quinnlue /tau-urban-acoustic-scenes-2022-mobile-split5 TAU Urban Acoustic Scenes 2022 Mobile - Split5 (Docs-aligned) This dataset repo contains a docs-aligned low-resource subset for DCASE-style evaluation: audio/: extracted audio files referenced by split5 + official test split evaluation_setup/fold1_train.csv: split5 training list (tab-separated, labeled) evaluation_setup/fold1_test.csv: test list (tab-separated, unlabeled) evaluation_setup/fold1_evaluate.csv: test list with labels (tab-separated) Source: Zenodo record 6337421 (TAU… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/tau-urban-acoustic-scenes-2022-mobile-split5.audio0 likes316 downloads5mo agoHugging Face11sanchit-gandhi /earnings22_splitWe partition the earnings22 dataset at https://huggingface.co/datasets/anton-l/earnings22_baseline_5_gram by source_id: Validation: 4420696 4448760 4461799 4469836 4473238 4482110 Test: 4432298 4450488 4470290 4479741 4483338 4485244 Train: remainder Official script for processing these splits will be released shortly. audio10K<n<100K0 likes295 downloads4y agoHugging Face12cairocode /MSPP_WAV_speaker_splitaudio100K<n<1M0 likes287 downloads1y agoHugging Face13malagasy-asr /malagasy-asr-100h-split Malagasy ASR 100h Split This dataset contains the cleaned and reviewed Malagasy ASR 100h split. Splits Split Samples Hours train 33505 90.1227 validation 1860 4.9134 test 1860 4.9640 Columns audio: relative path to the audio file text: reviewed reference transcription duration_sec: audio duration in seconds source: source label (waxal or voxlingua) Notes This split includes the final human review corrections… See the full description on the dataset page: https://huggingface.co/datasets/malagasy-asr/malagasy-asr-100h-split.audioautomatic-speech-recognition10K<n<100K0 likes249 downloads3d agoHugging Face14datTrantien17 /vlsp2020_vinbig_100h_splitaudio10K<n<100K0 likes221 downloads1y agoHugging Face15KasuleTrevor /sh_val_splitaudio100K<n<1M0 likes220 downloads2y agoHugging Face16bootcamp-labs /faceit_top_demos_883_voice_split_cut_20251126audio100K<n<1M0 likes220 downloads10mo agoHugging Face17keeve101 /common-voice-21.0-2025-03-14-zh-CN-splitaudio10K<n<100K1 likes206 downloads2y agoHugging Face18ghananlpcommunity /ghana-female-twi-8sec-splits This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 8-Word Speech Segments 25951 speech-text pairs split from 30-min recordings. Processing pipeline Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-female-twi-8sec-splits.audioautomatic-speech-recognition10K<n<100K0 likes192 downloads3mo agoHugging Face19rmndrnts /lena_dataset_splittedSubset-splitted version of a synthetic dataset based on specific domains. The textual basis was picked up by Elena Bruches. The non-splitted version can be found here: https://huggingface.co/datasets/rmndrnts/lena_dataset audio10K<n<100K0 likes187 downloads2y agoHugging Face20sanchit-gandhi /earnings22_split_resampledWe partition the earnings22 dataset at https://huggingface.co/datasets/anton-l/earnings22_baseline_5_gram by source_id: Validation: 4420696 4448760 4461799 4469836 4473238 4482110 Test: 4432298 4450488 4470290 4479741 4483338 4485244 Train: remainder Official script for processing these splits will be released shortly. audio10K<n<100K0 likes180 downloads4y agoHugging Face21ghananlpcommunity /ghana-female-twi-asr-16word-splits This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 25951 speech-text pairs split from 30-min recordings. Processing pipeline Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-female-twi-asr-16word-splits.audioautomatic-speech-recognition10K<n<100K0 likes173 downloads3mo agoHugging Face22Milana /resampled_16KHrz_vctk_speakers_splitaudio10K<n<100K0 likes159 downloads2y agoHugging Face23yigagilbert /luganda-english-cleaned-v1-splitgatedaudio100K<n<1M0 likes137 downloads5mo agoHugging Face24Ar4ikov /iemocap_audio_text_splitted Dataset Card for "iemocap_audio_text_splitted" More Information needed audio10K<n<100K3 likes133 downloads3y agoHugging Face25EYEDOL /naija-voices-hausa-split_0-1audio10K<n<100K0 likes131 downloads1y agoHugging Face26Sebssihakim /Clean_One_Speaker_STT_Split_EN_AR Clean_One_Speaker_STT_Split_EN_AR Mixed STT dataset — Saudi dialectal Arabic + Modern Standard Arabic + English — with train/validation/test splits, built for fine-tuning multilingual ASR (e.g. Whisper) without catastrophic forgetting of English. Composition Dialectal Arabic (~72%) — Sebssihakim/Clean_One_Speaker_SADA_Split (cleaned single-speaker SADA; SDAIA / Saudi Broadcasting Authority, CC BY-NC-SA 4.0) MSA (~8%) — FLEURS ar_eg, full official splits (CC-BY;… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_STT_Split_EN_AR.audioautomatic-speech-recognition10K<n<100K0 likes127 downloads3mo agoHugging Face27NgQuocThai /S2T_Split_NoRom_phase2audio10K<n<100K0 likes123 downloads11mo agoHugging Face28mmonzel /uwb_atcosim_split_by_person_6_2_2 UWB-ATCC + ATCOSIM — combined ATC ASR dataset (ATCOSIM 6-2-2, gender-balanced) A combined air-traffic-control ASR dataset built from Jzuluaga/uwb_atcc and Jzuluaga/atcosim_corpus, re-split with leakage-free (group-disjoint) train/validation/test splits. UWB-ATCC is split recording-disjoint (80/10/10) and ATCOSIM is split speaker-disjoint and gender-balanced, putting 1 male + 1 female speaker in each of validation and test (6 train / 2 / 2 speakers), with a fixed seed (42) and 16… See the full description on the dataset page: https://huggingface.co/datasets/mmonzel/uwb_atcosim_split_by_person_6_2_2.audio10K<n<100K1 likes118 downloads3mo agoHugging Face29EYEDOL /naija-voices-igbo-split_1-1audio10K<n<100K0 likes113 downloads1y agoHugging Face30EYEDOL /naija-voices-yoruba-split_2-3audio10K<n<100K0 likes110 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.