datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bolAIndia
bolAIndia
Human-side speech from production call recordings, cut into utterance-level
chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR
providers. Each row keeps the transcript, the provider's confidence, and full
provenance back to the source recording.
Sources
One config per transcription system, so their output stays separable.
config (source_id)
provider
model
hours
rows
shards
vendor-a
vendor-a
undisclosed
420.03
480774… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.rixvox-v2
RixVox-v2: A Swedish parliamentary speech dataset
RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.maa
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/khursani8/maa.central-kurdish-pseudolabel
Central Kurdish → English Pseudo-Labeled Speech Translation Corpus
Dataset Summary
This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish).
The dataset was automatically generated using a pipeline composed of:
Speech segmentation
Automatic Speech Recognition (ASR)
Machine Translation (MT)
The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.kasem-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Kasem Speech-Text Parallel Dataset
Dataset Description
This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.Kamba-ASR-Data-Subset-484H
Kamba ASR Data Subset 484H
Kamba speech dataset for automatic speech recognition.
ISSAI_KSC_335RS_v_1_1
Dataset Card for "ISSAI_KSC_335RS_v_1_1"
Kazakh Speech Corpus (KSC)
Identifier: SLR102
Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours)
Category: Speech
License: Attribution 4.0 International (CC BY 4.0)
Downloads (use a mirror closer to you):
ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN]
About this resource:
A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.spgispeech
Dataset Card for SPGISpeech
Dataset Description
Dataset Summary
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Annotations
Personal and Sensitive Information
Considerations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Citation Information
Contributions
Terms of Usage… See the full description on the dataset page: https://huggingface.co/datasets/kensho/spgispeech.SPGISpeech2.0
Dataset Card for SPGISpeech 2.0
Dataset Details
Dataset Overview
We are excited to present SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR).… See the full description on the dataset page: https://huggingface.co/datasets/kensho/SPGISpeech2.0.kiraat
KIRAAT — A Turkish Read-Speech Corpus
A sentence-aligned read-speech corpus built from publicly available
recordings on Turkish audiobook YouTube channels. The channel credits are
in the table at the end of this card; every clip carries the channel it came
from in the channel column.
clips
1,840,404
duration
3,105.7 hours
recommended subset
1,547,494 clips / 2,575.2 hours
channels
27
speakers (clustered)
90
source recordings
2,680
words (ASR)
21,695,774… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/kiraat.propagator-multimodal-pretraining-data
Propagator Multimodal Pretraining Data
This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format.
This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout.
Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.YO-CPT-kk
YO-CPT-kk
YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily
quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker,
TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a
punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and
cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the
voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.khm-asr-cultural
Khmer ASR Cultural Dataset
134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/khmer-speech-dataset.kupe-asr-en-data
kupe-asr-en-mini-150m — data
Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly):
raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this.
mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this.
Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state.
from datasets import load_dataset
ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train")
SUBAK.KO
Dataset Card for SUBAK.KO
Dataset Summary
SUBAK.KO (সুবাক্য), a publicly available annotated Bangladeshi standard Bangla speech corpus, is compiled for automatic speech recognition research.
This corpus contains 241 hours of high-quality speech data, including 229 hours of read speech data and 12 hours of broadcast speech data.
The read speech segment is recorded in a noise-proof studio environment from 33 male and 28 female native Bangladeshi Bangla speakers… See the full description on the dataset page: https://huggingface.co/datasets/SUST-CSE-Speech/SUBAK.KO.zeroth-korean
Zeroth-Korean
Zeroth-Korean
The data set contains transcriebed audio data for Korean. There are 51.6 hours transcribed Korean audio for training data (22,263 utterances, 105 people, 3000 sentences) and 1.2 hours transcribed Korean audio for testing data (457 utterances, 10 people). This corpus also contains pre-trained/designed language model, lexicon and morpheme-based segmenter(morfessor).
Zeroth project introduces free Korean speech corpus and aims to make Korean… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/zeroth-korean.nst
NST Swedish ASR Database (16 kHz) – reorganized
This database was created by Nordic Language Technology for the development of automatic speech recognition and dictation in Swedish. In this updated version, the organization of the data have been altered to improve the usefulness of the database.
In the original version of the material, the files were organized in a specific folder structure where the folder names were meaningful. However, the file names were not meaningful, and… See the full description on the dataset page: https://huggingface.co/datasets/KTH/nst.seamless-interaction-jefferson-annotations
Seamless Interaction Jefferson-Style Annotations
An automatic, turn-oriented annotation layer for the
Meta Seamless Interaction Dataset.
It compares the dataset's traditional transcript with an ASR-derived
Jefferson-style condition and supplies speech-act, communicative-purpose,
interactional-signal, alignment, and quality fields.
This is a derived noncommercial research dataset. It does not redistribute
the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.korean_datasetomnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/KathleenKunLiu/omnilingual-asr-corpus.kumawood-speech-transcriptions
Kumawood Speech Transcriptions
Speech segments from Ghanaian films, each paired with the film's human-authored
English subtitle and a machine Twi transcript.
Total number of hours 249.8 hours
Fields
field
meaning
audio
16 kHz mono FLAC segment
text
English subtitle displayed during the segment (human-authored, recovered by OCR)
twi_text
Twi transcript from Google STT (ak) — machine output
twi_words_per_sec
transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.KsponSpeechVartalaap
Vartalaap — full-duplex Hindi/English conversational speech
Dual-channel synthetic Indian customer-support calls for training full-duplex
speech-to-speech models.
68,674 calls · 1,564.4 hours · 1786 shards
(last updated 2026-09-24 10:34 IST)
Audio layout
Each row's audio is a stereo FLAC at 24000 Hz:
channel
content
0 (LEFT)
agent — pristine, TTS speech and silence only
1 (RIGHT)
user — the caller
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/Vartalaap.Kurisu_Voice
Makise Kurisu Multilingual Voice Dataset
13,999 labelled clips (17.1423 hours) of Makise Kurisu,
including the Amadeus Kurisu variant, in six languages, cut from the
STEINS;GATE games, the anime, and three character songs.
Every clip carries the transcript, a measured acoustic profile, an independent
speaker-identity check, and — where it could be earned rather than guessed — an
expressive tag. This is an unofficial, fan-made dataset with no affiliation to
the STEINS;GATE rights… See the full description on the dataset page: https://huggingface.co/datasets/starrydark/Kurisu_Voice.kumawood-speech-transcriptions
Kumawood Speech Transcriptions
Speech segments from Ghanaian films, each paired with the film's human-authored
English subtitle and a machine Twi transcript.
Total number of hours 249.8 hours
Fields
field
meaning
audio
16 kHz mono FLAC segment
text
English subtitle displayed during the segment (human-authored, recovered by OCR)
twi_text
Twi transcript from Google STT (ak) — machine output
twi_words_per_sec
transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.
