CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kapturecx /bolAIndiagated bolAIndia Human-side speech from production call recordings, cut into utterance-level chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR providers. Each row keeps the transcript, the provider's confidence, and full provenance back to the source recording. Sources One config per transcription system, so their output stays separable. config (source_id) provider model hours rows shards vendor-a vendor-a undisclosed 420.03 480774… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.audioautomatic-speech-recognition10M<n<100M1 likes12k downloads17m agoHugging Face02KBLab /rixvox-v2 RixVox-v2: A Swedish parliamentary speech dataset RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.audioautomatic-speech-recognition1M<n<10M12 likes6.7k downloads1y agoHugging Face03khursani8 /maa Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/khursani8/maa.audioautomatic-speech-recognition10M<n<100M7 likes5.7k downloads7mo agoHugging Face04aranemini /central-kurdish-pseudolabel Central Kurdish → English Pseudo-Labeled Speech Translation Corpus Dataset Summary This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish). The dataset was automatically generated using a pipeline composed of: Speech segmentation Automatic Speech Recognition (ASR) Machine Translation (MT) The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.audioautomatic-speech-recognition1M<n<10M2 likes4.2k downloads3mo agoHugging Face05Digital-Divide-Data /khmer-speech-dataset Khmer ASR Cultural Dataset 727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.audioautomatic-speech-recognition100K<n<1M26 likes3.8k downloads3mo agoHugging Face06ksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes3.4k downloads16d agoHugging Face07Ken-Z /Latin-Audio Dataset Summary Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training. Alignment and curation: Kaiyuan Zhao Language: Latin (Classical) Uses This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.audiotext-to-speech10K<n<100K8 likes3.2k downloads2mo agoHugging Face08ghanaopenai /kasem-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Kasem Speech-Text Parallel Dataset Dataset Description This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes2.9k downloads3mo agoHugging Face09Digital-Divide-Data /Kamba-ASR-Data-Subset-484H Kamba ASR Data Subset 484H Kamba speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M0 likes2.8k downloads1mo agoHugging Face10Shirali /ISSAI_KSC_335RS_v_1_1 Dataset Card for "ISSAI_KSC_335RS_v_1_1" Kazakh Speech Corpus (KSC) Identifier: SLR102 Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours) Category: Speech License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN] About this resource: A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.audioautomatic-speech-recognition100K<n<1M3 likes2.4k downloads4y agoHugging Face11kensho /spgispeechgated Dataset Card for SPGISpeech Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Terms of Usage… See the full description on the dataset page: https://huggingface.co/datasets/kensho/spgispeech.audioautomatic-speech-recognition1M<n<10M42 likes2.2k downloads8mo agoHugging Face12kensho /SPGISpeech2.0gated Dataset Card for SPGISpeech 2.0 Dataset Details Dataset Overview We are excited to present SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR).… See the full description on the dataset page: https://huggingface.co/datasets/kensho/SPGISpeech2.0.audioautomatic-speech-recognition100K<n<1M4 likes1.9k downloads5mo agoHugging Face13serdarcaglar /kiraatgated KIRAAT — A Turkish Read-Speech Corpus A sentence-aligned read-speech corpus built from publicly available recordings on Turkish audiobook YouTube channels. The channel credits are in the table at the end of this card; every clip carries the channel it came from in the channel column. clips 1,840,404 duration 3,105.7 hours recommended subset 1,547,494 clips / 2,575.2 hours channels 27 speakers (clustered) 90 source recordings 2,680 words (ASR) 21,695,774… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/kiraat.audiotext-to-speech1M<n<10M11 likes1.8k downloads12d agoHugging Face14ken-sungmin /propagator-multimodal-pretraining-data Propagator Multimodal Pretraining Data This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format. This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout. Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.texttext-generation0 likes1.3k downloads3mo agoHugging Face15humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.3k downloads8mo agoHugging Face16NCSpeech /YO-CPT-kk YO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.audiotext-to-speech100K<n<1M10 likes1.3k downloads2mo agoHugging Face17Digital-Divide-Data /khm-asr-cultural Khmer ASR Cultural Dataset 134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.audioautomatic-speech-recognition10K<n<100K9 likes1.3k downloads5mo agoHugging Face18phonsobon /khmer-speech-dataset Khmer ASR Cultural Dataset 727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/khmer-speech-dataset.audioautomatic-speech-recognition100K<n<1M0 likes1.3k downloads2mo agoHugging Face19anuj-inavlabs /kupe-asr-en-data kupe-asr-en-mini-150m — data Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly): raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this. mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this. Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state. from datasets import load_dataset ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train") tabularautomatic-speech-recognition1M<n<10M0 likes1k downloads15d agoHugging Face20SUST-CSE-Speech /SUBAK.KO Dataset Card for SUBAK.KO Dataset Summary SUBAK.KO (সুবাক্য), a publicly available annotated Bangladeshi standard Bangla speech corpus, is compiled for automatic speech recognition research. This corpus contains 241 hours of high-quality speech data, including 229 hours of read speech data and 12 hours of broadcast speech data. The read speech segment is recorded in a noise-proof studio environment from 33 male and 28 female native Bangladeshi Bangla speakers… See the full description on the dataset page: https://huggingface.co/datasets/SUST-CSE-Speech/SUBAK.KO.audioautomatic-speech-recognition10K<n<100K11 likes947 downloads3y agoHugging Face21Bingsu /zeroth-korean Zeroth-Korean Zeroth-Korean The data set contains transcriebed audio data for Korean. There are 51.6 hours transcribed Korean audio for training data (22,263 utterances, 105 people, 3000 sentences) and 1.2 hours transcribed Korean audio for testing data (457 utterances, 10 people). This corpus also contains pre-trained/designed language model, lexicon and morpheme-based segmenter(morfessor). Zeroth project introduces free Korean speech corpus and aims to make Korean… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/zeroth-korean.audioautomatic-speech-recognition10K<n<100K46 likes942 downloads4y agoHugging Face22KTH /nst NST Swedish ASR Database (16 kHz) – reorganized This database was created by Nordic Language Technology for the development of automatic speech recognition and dictation in Swedish. In this updated version, the organization of the data have been altered to improve the usefulness of the database. In the original version of the material, the files were organized in a specific folder structure where the folder names were meaningful. However, the file names were not meaningful, and… See the full description on the dataset page: https://huggingface.co/datasets/KTH/nst.audioautomatic-speech-recognition100K<n<1M4 likes888 downloads6mo agoHugging Face23kennethli319 /seamless-interaction-jefferson-annotations Seamless Interaction Jefferson-Style Annotations An automatic, turn-oriented annotation layer for the Meta Seamless Interaction Dataset. It compares the dataset's traditional transcript with an ASR-derived Jefferson-style condition and supplies speech-act, communicative-purpose, interactional-signal, alignment, and quality fields. This is a derived noncommercial research dataset. It does not redistribute the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.tabularautomatic-speech-recognition100K<n<1M0 likes873 downloads2mo agoHugging Face24JaepaX /korean_datasetaudioautomatic-speech-recognition10K<n<100K4 likes805 downloads2y agoHugging Face25KathleenKunLiu /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/KathleenKunLiu/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M0 likes775 downloads3mo agoHugging Face26ghananlpcommunity /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes758 downloads22d agoHugging Face27jp1924 /KsponSpeechgatedaudioautomatic-speech-recognition100K<n<1M12 likes737 downloads9mo agoHugging Face28kapturecx /Vartalaapgated Vartalaap — full-duplex Hindi/English conversational speech Dual-channel synthetic Indian customer-support calls for training full-duplex speech-to-speech models. 68,674 calls · 1,564.4 hours · 1786 shards (last updated 2026-09-24 10:34 IST) Audio layout Each row's audio is a stereo FLAC at 24000 Hz: channel content 0 (LEFT) agent — pristine, TTS speech and silence only 1 (RIGHT) user — the caller from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/Vartalaap.audioautomatic-speech-recognition10K<n<100K0 likes643 downloads23h agoHugging Face29starrydark /Kurisu_Voice Makise Kurisu Multilingual Voice Dataset 13,999 labelled clips (17.1423 hours) of Makise Kurisu, including the Amadeus Kurisu variant, in six languages, cut from the STEINS;GATE games, the anime, and three character songs. Every clip carries the transcript, a measured acoustic profile, an independent speaker-identity check, and — where it could be earned rather than guessed — an expressive tag. This is an unofficial, fan-made dataset with no affiliation to the STEINS;GATE rights… See the full description on the dataset page: https://huggingface.co/datasets/starrydark/Kurisu_Voice.audiotext-to-speech10K<n<100K0 likes609 downloads20d agoHugging Face30ghanaopenai /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes596 downloads22d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.