datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.ISSAI_KSC_335RS_v_1_1
Dataset Card for "ISSAI_KSC_335RS_v_1_1"
Kazakh Speech Corpus (KSC)
Identifier: SLR102
Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours)
Category: Speech
License: Attribution 4.0 International (CC BY 4.0)
Downloads (use a mirror closer to you):
ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN]
About this resource:
A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.KsponSpeechmultiturn_ks
khursanirevo/multiturn_ks
Dataset Description
Multiturn dialogue dataset with speaker-separated stereo audio and multi-language transcripts from 139 YouTube videos.
Features
Audio: Stereo audio with speaker separation (speaker 0 = left channel, speaker 1 = right channel)
Segments: Speaker turn-level annotations with timestamps for English and Malay
Multi-language: Transcripts in 9 languages (en, ms, zh-Hans, zh-Hant, ru, id, ar, ja, ko)
Video ID: YouTube video… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/multiturn_ks.superb_ksThe Superb dataset for the Keyword Spotting (KS) task without needing to run remote code, so it is compatible with datasets >= 4.0.0.
ukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples
validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.kazakh_speech_dataset_ksdKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api.
Dataset info:
813 Speakers
with 500 samples for 4 speakers
with 250 samples for 809 speakers
Male/female
555 Hours
Guides
Load data 1
Replace the export HF_HOME with your HF_HOME path
from datasets import load_dataset
# export HF_HOME="/data/vladimir_albrekht/hf_cache"
ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.kyrgyz-asr-verified-v1🇰🇬 Kyrgyz ASR-Verified Speech Corpus
metric
value
Clips
411,834
Audio
874 hours
Speakers
280
Verification CER
mean 0.79%, median 0.00%, p90 2.35%
License & access
KSTU members only. This dataset is released under a custom KSTU Internal Dataset License (other, see the LICENSE file):
🎓 Access and use are restricted to KSTU (Kyrgyz State Technical University) members and KSTU-affiliated persons.
🚫 Requests from outside KSTU will not be approved. Request… See the full description on the dataset page: https://huggingface.co/datasets/kstunlp/kyrgyz-asr-verified-v1.
