CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M489 likes48k downloads2y agoHugging Face02SparkAudio /voxbox VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ # Audio files (organised by sub-corpus) │ └── ... └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice_cn.jsonl ├── ... └── wenetspeech4tts.jsonl # JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.audiotext-to-speech10M<n<100M76 likes45k downloads1y agoHugging Face03sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes33k downloads10mo agoHugging Face04laion /LAION-Audio-300Maudio100M<n<1B74 likes18k downloads2y agoHugging Face05laion /soundscapesaudio10M<n<100M7 likes16k downloads1y agoHugging Face06qualialabsAI /DuplexConv DuplexConv DuplexConv is a large-scale Chinese multi-channel conversational speech dataset with LLM-assisted annotations, developed by ASLP@NPU and QualiaLabs as part of the SmoothConv–DuplexConv corpus family. Companion dataset: SmoothConv on HuggingFace (100 hours, expert human annotation). DuplexConv and SmoothConv share the same conversational domains and a unified data design. SmoothConv focuses on high-quality human annotations for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/qualialabsAI/DuplexConv.audio100K<n<1M13 likes16k downloads3mo agoHugging Face07laion /laions_got_talent LAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel. Dataset Composition The dataset includes: Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.audio100K<n<1M41 likes9.8k downloads2y agoHugging Face08sarulab-speech /mls_sidon MLS-Sidon Overview This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. The dataset is provided in WebDataset format for efficient large-scale training. Source: Multilingual LibriSpeech Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese Format: WebDataset (.tar shards) License: CC-BY-4.0 Dataset Structure Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.audiotext-to-speech10M<n<100M11 likes7.9k downloads1y agoHugging Face09speechcolab /gigaspeech2gated Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.audioautomatic-speech-recognition10M<n<100M71 likes6k downloads6mo agoHugging Face10XRXRX /X-Voice-Dataset-Train X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling. Also the train set of X-Voice Model. Core Statistics Total Speech Duration: 420K hours 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.audiotext-to-speech10M<n<100M11 likes4.8k downloads5mo agoHugging Face11laion /laion-audio-previewaudio1M<n<10M11 likes4.5k downloads2y agoHugging Face12osanpo /reazonspeechaudio10K<n<100K2 likes4.4k downloads3y agoHugging Face13mitermix /audiosnippetsaudio1M<n<10M6 likes4.3k downloads2y agoHugging Face14Vyvo-Research /Emilia-YODAS-ENaudio10M<n<100M4 likes3.8k downloads10mo agoHugging Face15lighthouse-emnlp2024 /Clotho-Moment Clotho-Moment This repository provides wav files used in Language-based Audio Moment Retrieval. Each sample includes long audio containing some audio events with the temporal and textual annotation. Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/ Code: https://github.com/line/lighthouse Split Train train/train-{000..715}.tar 37930 audio samples Valid valid/valid-{000..108}.tar 5741 audio samples Test test/test-{000..142}.tar 7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.audioaudio-text-to-text10K<n<100K2 likes3.7k downloads8mo agoHugging Face16laion /laions_got_talent_rawaudio10K<n<100K7 likes3.5k downloads2y agoHugging Face17TalTechNLP /voxlingua107_wds VoxLingua107 VoxLingua107 is a speech dataset for training spoken language identification models. The dataset consists of short speech segments automatically extracted from YouTube videos and labeled according the language of the video title and description, with some post-processing steps to filter out false positives. VoxLingua107 contains data for 107 languages. The total amount of speech in the training set is 6628 hours. The average amount of data per language is 62 hours.… See the full description on the dataset page: https://huggingface.co/datasets/TalTechNLP/voxlingua107_wds.audio1M<n<10M4 likes3.4k downloads1y agoHugging Face18TTS-AGI /emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset. https://huggingface.co/datasets/amphion/Emilia-Dataset audiotext-to-speech10M<n<100M5 likes3.1k downloads2y agoHugging Face19mitermix /audiosnippets_small_with_detailed_annotationaudio100K<n<1M1 likes2.7k downloads2y agoHugging Face20mitermix /audiosnippets_small_with_detailed_annotation2audio1M<n<10M1 likes2.7k downloads2y agoHugging Face21bhyuan /gptsovits_datasetgated bhyuan/gptsovits_dataset GPT-SoVITS speech dataset, packed as WebDataset tar shards. Layout data/ train/ metadata.csv audio/ train-000.tar train-001.tar ... validation/ metadata.csv audio/ validation-000.tar ... test/ metadata.csv audio/ test-000.tar ... Shard counts: youshengshu_v5_test: 6536 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.audiotext-to-speech10M<n<100M1 likes2.5k downloads4mo agoHugging Face22zh-liu799 /0813audio100K<n<1M0 likes2.4k downloads1y agoHugging Face23TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes2.4k downloads6mo agoHugging Face24laion /Emolia Dataset Card for Emolia Dataset Description This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?". The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.audio10M<n<100M15 likes2.2k downloads10mo agoHugging Face25laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes2.1k downloads11mo agoHugging Face26laion /majestrino-dataaudio1M<n<10M1 likes1.9k downloads6mo agoHugging Face27THUdyh /Ola-DataThis repository contains the data presented in Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment. Code: https://github.com/Ola-Omni/Ola audioany-to-any100K<n<1M9 likes1.7k downloads2y agoHugging Face28gijs /voice-data Voice-Data: a curated multi-corpus voice dataset for voice–text contrastive training voice-data is a single, globally-shuffled WebDataset that bundles several voice/speech corpora into one ready-to-train mixture for voice–text contrastive (CLAP-style) models such as VoiceCLAP. Each clip pairs 48 kHz mono FLAC audio with a natural-language text caption describing the voice — its emotion, prosody, timbre, speaking style, recording context, and speaker traits. The distinguishing… See the full description on the dataset page: https://huggingface.co/datasets/gijs/voice-data.audioaudio-classification1M<n<10M0 likes1.7k downloads4mo agoHugging Face29yhytoto12 /behavior-sd 🎙️ Behavior-SD Official repository for our NAACL 2025 paper:Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language ModelsSehun Lee*, Kang-wook Kim*, Gunhee Kim (* Equal contribution) 🏆 SAC Award Winner in Speech Processing and Spoken Language Understanding 🔗 Links Project Page Code 📖 Overview We explores how to generate natural, behaviorally-rich full-duplex spoken dialogues using large language models (LLMs). We introduce:… See the full description on the dataset page: https://huggingface.co/datasets/yhytoto12/behavior-sd.audio100K<n<1M9 likes1.6k downloads1y agoHugging Face30krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.