CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shangeth /expresso Expresso — audio + text A faithful re-publication of the official Expresso dataset (Nguyen et al., Interspeech 2023) as a loadable HuggingFace audio dataset, sourced directly from FAIR's official tar. ⚠️ License: CC-BY-NC-4.0 — non-commercial use only. Configs read — 11.6k mono read-speech utterances with human transcripts. conversational — ~15.9k mono per-utterance turns derived from the stereo conversational dialogues, transcribed with Whisper Large V3 Turbo.… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/expresso.audiotext-to-speech10K<n<100K0 likes245 downloads5mo agoHugging Face02shannonnonshan /ViMedCSS-Cop 🩺 ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset (LREC 2026) 📖 Overview ViMedCSS is a Vietnamese medical speech dataset for code-switching ASR, where each utterance contains at least one non-Vietnamese (mainly English) medical term embedded in Vietnamese speech. 📊 Dataset Statistics Split Statistics (from ViMedCSS-Metadata) Split # Rows Duration (hours) Avg duration (s) Total CS terms train 11,832 24.30 7.39 12,314… See the full description on the dataset page: https://huggingface.co/datasets/shannonnonshan/ViMedCSS-Cop.audioautomatic-speech-recognition10K<n<100K1 likes203 downloads7mo agoHugging Face03freococo /shan_language_asr_voices ⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy. This archive stands as one of the largest publicly… See the full description on the dataset page: https://huggingface.co/datasets/freococo/shan_language_asr_voices.audioautomatic-speech-recognition10K<n<100K2 likes93 downloads1y agoHugging Face04shangeth /ljspeech-mimi-codes LJSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LJSpeech corpus — 13,100 English utterances from a single female speaker reading public-domain audiobook passages (~24 hours). This dataset contains codes only, not audio. For waveforms, go to the original LJSpeech release; these codes are designed to be loaded alongside it for training Mimi-based speech models without paying the ~1 hour of GPU extraction cost. Schema One row per utterance:… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/ljspeech-mimi-codes.tabulartext-to-speech10K<n<100K0 likes71 downloads5mo agoHugging Face05shangeth /librispeech-mimi-codes LibriSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LibriSpeech corpus — multi-speaker English audiobook readings from the LibriVox project. This dataset contains codes only, not audio. For waveforms, use any of the LibriSpeech mirrors (e.g. openslr/librispeech_asr); these codes let you skip the ~hours of GPU extraction needed to train Mimi-based speech models. Schema One row per utterance: Column Type Notes id string… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/librispeech-mimi-codes.tabulartext-to-speech100K<n<1M0 likes55 downloads5mo agoHugging Face06shangeth /mls-mimi-codes Multilingual LibriSpeech (MLS) — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for Multilingual LibriSpeech — LibriVox audiobooks in 7 non-English languages. English is intentionally excluded. For English Mimi codes, use: shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits) shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native) shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents) shangeth/jenny-mimi-codes — Jenny… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.tabulartext-to-speech1M<n<10M0 likes48 downloads5mo agoHugging Face07shangeth /expresso-mimi-codes Expresso — Mimi Codes (k = 32) Pre-extracted Kyutai Mimi tokens (all 32 codebooks) for both the read and conversational subsets of Expresso. Source audio + transcripts live in shangeth/expresso; this dataset publishes the discrete-token version for training Mimi-based speech models without re-extracting. ⚠️ License: CC-BY-NC-4.0 — non-commercial use only. Why Expresso for Wren? Expresso is the most directly relevant dataset for speech disentanglement research — the… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/expresso-mimi-codes.tabulartext-to-speech10K<n<100K1 likes35 downloads5mo agoHugging Face08freococo /rfa_shan_language_voices RFA Shan Language Voices This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_shan_language_voices.audioautomatic-speech-recognition1K<n<10K0 likes30 downloads1y agoHugging Face09shane062 /FYP_Fine_Tuning Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/shane062/FYP_Fine_Tuning.audioautomatic-speech-recognitionn<1K1 likes29 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.