CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shraavb /spanish-slang-stt-data Spanish Regional Speech-to-Text Dataset A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models. Dataset Description This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions: Region Samples Description Mexico 17,725 Mexican Spanish including CIEMPIESS corpus Spain 11,360 Castilian Spanish from TEDx and Common Voice Argentina 5,839 Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.audioautomatic-speech-recognition10K<n<100K0 likes895 downloads8mo agoHugging Face02rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes266 downloads1y agoHugging Face03Codyfederer /tr-full-dataset TR-Full_dataset This is a merged speech dataset containing 41427 audio segments from 88 source datasets. Dataset Information Total Segments: 41427 Speakers: 222 Languages: tr Emotions: neutral, angry, sad, happy Original Datasets: 88 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.audioautomatic-speech-recognition10K<n<100K6 likes145 downloads1y agoHugging Face04Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes138 downloads1mo agoHugging Face05rorosese /my-voxtral-datasetaudion<1K0 likes135 downloads1y agoHugging Face06rodenhhh /ContextTTS_dataset ContextTTS Evaluation Dataset This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks. Dataset Summary The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.audiotext-to-speechn<1K0 likes113 downloads6mo agoHugging Face07ZHANGYUXUAN-zR /MCF-Dataset MCF: Text LLMS For Multimodal Emotional Causality Data Dataset task definition and annotation example of the MCF framework. The framework contains two core subtasks: five-tuple element extraction (identifying Target, Holder, Aspect, Opinion, Sentiment, and Rationale) and sentiment chain analysis (constructing causal relationship chains between emotional events). The dataset is provided with the following structure. Each sample includes video, audio, and dialogue… See the full description on the dataset page: https://huggingface.co/datasets/ZHANGYUXUAN-zR/MCF-Dataset.audiotext-classification10K<n<100K2 likes72 downloads1y agoHugging Face08demegire /personaplex-finetuning-pharma-data-sample PersonaPlex Finetuning — Pharma Data Sample A 10-example slice of the synthetic patient-support / medication adherence dataset used to train demegire/personaplex-finetune-pharma. The on-disk layout below is exactly what the trainer in emotion-machine-org/personaplex-finetune consumes — use this as a template when building your own. Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at sample scale). Layout . ├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.audiotext-to-speechn<1K0 likes66 downloads4mo agoHugging Face09David-A-Amoo /naijavoices_dataset_85_hours_tts_bestFull&nbsp;dataset&nbsp;re-upload&nbsp;with&nbsp;more&nbsp;statistics&nbsp;as&nbsp;well&nbsp;as&nbsp;filtering&nbsp;scripts&nbsp;that&nbsp;give&nbsp;the&nbsp;top&nbsp;x&nbsp;files&nbsp;or&nbsp;best&nbsp;x&nbsp;hours&nbsp;for&nbsp;tts&nbsp;based&nbsp;and&nbsp;calculations&nbsp;from&nbsp;acoustic&nbsp;metrics\color{Blue}{\large \textbf{Full dataset re-upload with more statistics as well as filtering scripts that give the top x files or best x hours for tts based and calculations from acoustic… See the full description on the dataset page: https://huggingface.co/datasets/David-A-Amoo/naijavoices_dataset_85_hours_tts_best.tabular10K<n<100K2 likes62 downloads2mo agoHugging Face10sunbv56 /song_dataset 🎵 Vietnamese Song Lyrics and Word Timestamps Dataset Dataset Summary The song_dataset provides high-quality Vietnamese song data, including metadata, full lyrics, and particularly word-level timestamps. This dataset is optimally designed for tasks such as: Training and evaluating automatic speech recognition (ASR) models on music. Lyrics synchronization (Lyrics Alignment / Karaoke generation). Natural language processing (NLP) analysis on song lyrics. The data is… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset.text1K<n<10K0 likes39 downloads7mo agoHugging Face11jeju-potato /jeju_potato_datasetsaudio10K<n<100K0 likes36 downloads1y agoHugging Face12electron-rare /mascarade-dsp-dataset Mascarade — DSP & Signal Processing Q&A ✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11) Per-sample Stack Exchange Electronics attribution recovered via the SE /search/advanced + /questions/{id} API search : 169 samples (~5.35 %) confirmed as Stack Exchange Electronics (CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution (URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60). 535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.texttext-generation1K<n<10K0 likes30 downloads5mo agoHugging Face13mira-iitjmu /ns-urdu-datasetaudio1K<n<10K0 likes29 downloads5mo agoHugging Face14ajey6789 /bhashini-datasetaudio1K<n<10K0 likes28 downloads1y agoHugging Face15galammadin-asr /chlid-datasetaudio10K<n<100K0 likes26 downloads7mo agoHugging Face16chtugha /small-german-medical-dialogue-dataset-for-moshi Small german dialogue dataset This dataset contains 500 completely made up medical phonecall dialogues between patients and a GP's office. Dataset Details Dataset Description 500 made up phonecalls that were first created with AI as text. The audio was then created using Openai tts-1-hd and the accurately timestamped transcripts were added. The audio files are formatted like this: Stereo with split channels: Speaker A is on the left channel… See the full description on the dataset page: https://huggingface.co/datasets/chtugha/small-german-medical-dialogue-dataset-for-moshi.audioaudio-text-to-textn<1K0 likes24 downloads4mo agoHugging Face17Letian2003 /stage1a_smoke_data stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh) Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP → frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend: WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT encoder — no raw-audio decoding at train time. 113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.tabularautomatic-speech-recognitionn<1K0 likes23 downloads2mo agoHugging Face18TumeloKonaite /synthetic-patient-dr-data Synthetic Patient DR Data Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio. Dataset Summary This dataset was generated for research and prototyping in: clinical dialogue generation structured clinical extraction text-to-audio workflows conversational healthcare modeling All consultations are synthetic and should not be treated as real clinical encounters. Export Metadata Mode: audio Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.audiotext-generationn<1K0 likes18 downloads6mo agoHugging Face19sunbv56 /song_dataset_chunked Vietnamese Songs Word-Level Timestamp Dataset (Chunked) This dataset contains word-level timestamp information for Vietnamese songs, specifically pre-chunked into segments up to 30 seconds for use in training or fine-tuning speech recognition (ASR) systems like Whisper. Dataset Summary The song_dataset_chunked provides high-quality Vietnamese song data, properly segmented into optimal ~30-second sequences. Duration Insights: Train split (train_chunked.jsonl): ~ 230.62… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset_chunked.tabular10K<n<100K0 likes17 downloads7mo agoHugging Face20yadorigi /Onomatopoeia_Dataset🎧 Onomatopoeia Dataset (Audio → Manga Expression) 音声解析結果をもとに、日本語のオノマトペ(擬音語・擬態語)を生成するためのデータセットです。 本データセットは、音そのものではなく、音から推定された特徴・空間・情景を入力とする構造化データであり、 漫画的な表現生成を目的としたマルチモーダルデータです。 📌 Dataset Summary 本データセットは以下のパイプラインから生成されています: Audio ↓ Audio Features (04_features.json) ↓ Audio Events (05_audio_events.json) ↓ Space Judgement (06_space_judgement.json) ↓ Scene Interpretation (07_scene_interpretation.json) ↓ Onomatopoeia (08_onomatopoeia.json) 👉 音 → 空間 → 情景 → オノマトペ という段階的生成構造を持ちます。 📊… See the full description on the dataset page: https://huggingface.co/datasets/yadorigi/Onomatopoeia_Dataset.texttext-generationn<1K0 likes16 downloads6mo agoHugging Face21mszawerd /concept-datasetaudio1K<n<10K0 likes13 downloads2y agoHugging Face22Tnaot /SPS-Bopha-Voice-Dataset-v1gated VibeVoice Fine-Tuning Dataset: SPS-Bopha-Voice-Dataset-v1 This dataset is formatted for fine-tuning VibeVoice. Structure training_data.jsonl: The main manifest file containing transcriptions and paths. chunks_staging/: Directory containing the audio clips. Usage with VibeVoice Clone this repository: git clone https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1 cd SPS-Bopha-Voice-Dataset-v1 Run the training script pointing to… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1.audiotext-to-speech1K<n<10K0 likes13 downloads10mo agoHugging Face23anian0707 /scasr_datasetaudio10K<n<100K0 likes13 downloads1mo agoHugging Face24Ailiance-fr /mascarade-dsp-dataset Mascarade — DSP & Signal Processing Q&A ✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11) Per-sample Stack Exchange Electronics attribution recovered via the SE /search/advanced + /questions/{id} API search : 169 samples (~5.35 %) confirmed as Stack Exchange Electronics (CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution (URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60). 535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.texttext-generation1K<n<10K0 likes12 downloads5mo agoHugging Face25solbon1212 /fma-dataset-aug-caption FMA-CLAP Caption Augmentation Dataset Overview This dataset is an enhanced version of the FMA (Free Music Archive) dataset, where we have augmented the original metadata with natural language captions generated using the CLAP (Contrastive Language-Audio Pretraining) model. The captions describe the genre, style, mood, and instrumentation of each track, making it more suitable for zero-shot learning, music classification, and text-to-music generation tasks.… See the full description on the dataset page: https://huggingface.co/datasets/solbon1212/fma-dataset-aug-caption.textaudio-classification1K<n<10K1 likes11 downloads2y agoHugging Face26RidheshBhati /Indic_New_dataset_TTS Indic TTS Dataset Hub (Mozilla) Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice. Select the language from the Subset dropdown in the Dataset Viewer. Columns audio: WAV audio clip (16kHz) text: transcription duration: length in seconds speaking_rate: characters per second audio10K<n<100K0 likes11 downloads7mo agoHugging Face27Sopho /sopho-poetry-tts-data sopho-poetry-tts-data Training data for the SuFei on-device Chinese-poetry TTS pipeline (sopho-poetry-tts-train). Two co-equal lineages, one per teacher model — both active (the FS2 set is the provenance of production v9 and stays reusable). Lineage Teacher Trained Poems cosyvoice3/ Fun-CosyVoice3-0.5B (Apache-2.0) m3_v6 1023 paddlespeech_fs2/ PaddleSpeech FS2 CSMSC (Apache-2.0) v9 320 Layout poems.jsonl # shared source text, keyed… See the full description on the dataset page: https://huggingface.co/datasets/Sopho/sopho-poetry-tts-data.audiotext-to-speech1K<n<10K0 likes10 downloads2mo agoHugging Face28chihhab999 /topipl-spanish-datasetaudio1K<n<10K0 likes9 downloads9mo agoHugging Face29IbraahimLab /voice-dataset Voice Dataset Collected from the web uploader tool. Voice Dataset Collected from the web uploader tool. audion<1K0 likes9 downloads7mo agoHugging Face30OpenDCAI /dataflow-mm-audio_asr_pipelineaudio1K<n<10K0 likes6 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.