datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.ISSAI_KSC_335RS_v_1_1
Dataset Card for "ISSAI_KSC_335RS_v_1_1"
Kazakh Speech Corpus (KSC)
Identifier: SLR102
Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours)
Category: Speech
License: Attribution 4.0 International (CC BY 4.0)
Downloads (use a mirror closer to you):
ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN]
About this resource:
A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.ksponspeechKSS_Dataset
Description of the original author
KSS Dataset: Korean Single speaker Speech Dataset
KSS Dataset is designed for the Korean text-to-speech task. It consists of audio files recorded by a professional female voice actoress and their aligned text extracted from my books. As a copyright holder, by courtesy of the publishers, I release this dataset to the public. To my best knowledge, this is the first publicly available speech dataset for Korean.
File Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/KSS_Dataset.KsponSpeechksponspeech_03multiturn_ks
khursanirevo/multiturn_ks
Dataset Description
Multiturn dialogue dataset with speaker-separated stereo audio and multi-language transcripts from 139 YouTube videos.
Features
Audio: Stereo audio with speaker separation (speaker 0 = left channel, speaker 1 = right channel)
Segments: Speaker turn-level annotations with timestamps for English and Malay
Multi-language: Transcripts in 9 languages (en, ms, zh-Hans, zh-Hant, ru, id, ar, ja, ko)
Video ID: YouTube video… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/multiturn_ks.ksponspeech_05ksponspeech_04superb_ksThe Superb dataset for the Keyword Spotting (KS) task without needing to run remote code, so it is compatible with datasets >= 4.0.0.
KsponSpeechksponspeech-evalpaper link: https://www.mdpi.com/846876
ukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples
validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.superb_ks_synthk_speechKsponspeech_timestampsKSC2superb_ks
Dataset Card for "superb_ks"
More Information needed
lavrova_kd_ruKsponSpeech
Dataset Card for KsponSpeech
Dataset Summary
The KsponSpeech is a large-scale spontaneous speech corpus in Korean. This corpus contains 969 hours of general open-domain dialog utterances, spoken by approximately 2,000 native Korean speakers in a clean environment. The data was collected by recording dialogues between two people conversing freely on various topics, and then manually transcribing the utterances.
Please note that we are only sharing the evaluation set of… See the full description on the dataset page: https://huggingface.co/datasets/jubang0219/KsponSpeech.ksd_120hours_kkKsponSpeech_eval_otherksponspeech_eval_cleanksa_arabic_speech_muhammadNepali_ASR_Dataksponspeech_eval_clean_testKsponSpeech-eval-cleanKSS
Disclaimer
This dataset is not mine and I do not accept any legal responsibility for its use. This dataset is simply being reuploaded for easier accessibility.
kazakh_speech_dataset_ksdKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api.
Dataset info:
813 Speakers
with 500 samples for 4 speakers
with 250 samples for 809 speakers
Male/female
555 Hours
Guides
Load data 1
Replace the export HF_HOME with your HF_HOME path
from datasets import load_dataset
# export HF_HOME="/data/vladimir_albrekht/hf_cache"
ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.k_whisper_dataset
