CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes2.6k downloads15d agoHugging Face02Shirali /ISSAI_KSC_335RS_v_1_1 Dataset Card for "ISSAI_KSC_335RS_v_1_1" Kazakh Speech Corpus (KSC) Identifier: SLR102 Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours) Category: Speech License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN] About this resource: A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.audioautomatic-speech-recognition100K<n<1M3 likes2k downloads4y agoHugging Face03DragonLine /ksponspeechaudio100K<n<1M2 likes940 downloads3y agoHugging Face04Bingsu /KSS_Dataset Description of the original author KSS Dataset: Korean Single speaker Speech Dataset KSS Dataset is designed for the Korean text-to-speech task. It consists of audio files recorded by a professional female voice actoress and their aligned text extracted from my books. As a copyright holder, by courtesy of the publishers, I release this dataset to the public. To my best knowledge, this is the first publicly available speech dataset for Korean. File Format Each… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/KSS_Dataset.audiotext-to-speech10K<n<100K20 likes869 downloads4y agoHugging Face05jp1924 /KsponSpeechgatedaudioautomatic-speech-recognition100K<n<1M12 likes722 downloads9mo agoHugging Face06DragonLine /ksponspeech_03audio100K<n<1M0 likes467 downloads3y agoHugging Face07khursanirevo /multiturn_ks khursanirevo/multiturn_ks Dataset Description Multiturn dialogue dataset with speaker-separated stereo audio and multi-language transcripts from 139 YouTube videos. Features Audio: Stereo audio with speaker separation (speaker 0 = left channel, speaker 1 = right channel) Segments: Speaker turn-level annotations with timestamps for English and Malay Multi-language: Transcripts in 9 languages (en, ms, zh-Hans, zh-Hant, ru, id, ar, ja, ko) Video ID: YouTube video… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/multiturn_ks.audioautomatic-speech-recognition10K<n<100K0 likes308 downloads5mo agoHugging Face08DragonLine /ksponspeech_05audio100K<n<1M0 likes300 downloads3y agoHugging Face09DragonLine /ksponspeech_04audio100K<n<1M0 likes297 downloads3y agoHugging Face10regisss /superb_ksThe Superb dataset for the Keyword Spotting (KS) task without needing to run remote code, so it is compatible with datasets >= 4.0.0. audioaudio-classification10K<n<100K0 likes127 downloads1y agoHugging Face11PThi35 /KsponSpeechaudio10K<n<100K0 likes103 downloads2mo agoHugging Face12yfyeung /ksponspeech-evalpaper link: https://www.mdpi.com/846876 audioautomatic-speech-recognition0 likes102 downloads2y agoHugging Face13KSE-RESEARCH-Group /ukr-dialects-audio-dataset Ukrainian Dialects Audio Dataset Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits. Dataset Description This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets: NaUKMA-Audio-Dataset Ivanna-Stefiuk-Audio-Dataset Larysa-Irodenko-Audio-Dataset Hutsulendia-Audio-Dataset Dido-Yvanchyk-Audio-Dataset-v2 Dataset Structure train: 27,675 samples validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.audioautomatic-speech-recognition10K<n<100K1 likes89 downloads7mo agoHugging Face14Codec-SUPERB /superb_ks_synthaudio10K<n<100K0 likes77 downloads3y agoHugging Face15gcasey2 /k_speechaudio10K<n<100K0 likes62 downloads3y agoHugging Face16PThi35 /Ksponspeech_timestampsaudio10K<n<100K0 likes55 downloads3mo agoHugging Face17diaslmb /KSC2audio100K<n<1M0 likes47 downloads4mo agoHugging Face18Codec-SUPERB /superb_ks Dataset Card for "superb_ks" More Information needed audio10K<n<100K0 likes45 downloads3y agoHugging Face19ksgraves /lavrova_kd_ruaudion<1K0 likes38 downloads10mo agoHugging Face20jubang0219 /KsponSpeech Dataset Card for KsponSpeech Dataset Summary The KsponSpeech is a large-scale spontaneous speech corpus in Korean. This corpus contains 969 hours of general open-domain dialog utterances, spoken by approximately 2,000 native Korean speakers in a clean environment. The data was collected by recording dialogues between two people conversing freely on various topics, and then manually transcribing the utterances. Please note that we are only sharing the evaluation set of… See the full description on the dataset page: https://huggingface.co/datasets/jubang0219/KsponSpeech.audio1K<n<10K2 likes34 downloads2y agoHugging Face21gassirbek /ksd_120hours_kkaudio10K<n<100K1 likes32 downloads2y agoHugging Face22Kyudan /KsponSpeech_eval_otheraudio1K<n<10K0 likes22 downloads11mo agoHugging Face23DragonLine /ksponspeech_eval_cleanaudio1K<n<10K0 likes19 downloads3y agoHugging Face24xtz999 /ksa_arabic_speech_muhammadaudio1K<n<10K0 likes19 downloads4mo agoHugging Face25kshitizzzzzzz /Nepali_ASR_Dataaudio1K<n<10K0 likes17 downloads2y agoHugging Face26DragonLine /ksponspeech_eval_clean_testaudio1K<n<10K0 likes16 downloads3y agoHugging Face27Kyudan /KsponSpeech-eval-cleanaudio1K<n<10K0 likes15 downloads11mo agoHugging Face28Razer112 /KSS Disclaimer This dataset is not mine and I do not accept any legal responsibility for its use. This dataset is simply being reuploaded for easier accessibility. audion<1K0 likes15 downloads2mo agoHugging Face29SRP-base-model-training /kazakh_speech_dataset_ksdgatedKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api. Dataset info: 813 Speakers with 500 samples for 4 speakers with 250 samples for 809 speakers Male/female 555 Hours Guides Load data 1 Replace the export HF_HOME with your HF_HOME path from datasets import load_dataset # export HF_HOME="/data/vladimir_albrekht/hf_cache" ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.audioautomatic-speech-recognition100K<n<1M2 likes14 downloads1y agoHugging Face30k-seungri /k_whisper_datasetaudion<1K0 likes13 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.