CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /Emolia Dataset Card for Emolia Dataset Description This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?". The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.audio10M<n<100M15 likes2.1k downloads10mo agoHugging Face02krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face03krishnakalyan3 /emo_parleraudio1M<n<10M2 likes1.4k downloads2y agoHugging Face04krishnakalyan3 /emo_webdsaudio10K<n<100K5 likes1.3k downloads2y agoHugging Face05laion /Emilia-with-Emotion-Annotations4audio10M<n<100M1 likes895 downloads1y agoHugging Face06akuzdeuov /qwen3-tts-multilingual-emotional-speechaudio1M<n<10M0 likes748 downloads14d agoHugging Face07laion /Emilia-with-Emotion-Annotations5audio10M<n<100M3 likes735 downloads1y agoHugging Face08krishnakalyan3 /emo_speech_filtered_v12 second filtered emotional speech in webdataset format https://huggingface.co/datasets/EQ4You/Emotional_Speech audio10K<n<100K0 likes634 downloads2y agoHugging Face09laion /Emilia-with-Emotion-Annotations3audio10M<n<100M1 likes383 downloads1y agoHugging Face10VoiceNet /emolia emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired) This is the emolia-balanced-5M-subset corpus repackaged for high-quality audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz (PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json samples. The JSON sidecar carries the full annotation stack: Original metadata (id, text, duration, speaker, language, dnsmos). A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.audioaudio-classification1M<n<10M1 likes381 downloads5mo agoHugging Face11laion /Emilia-with-Emotion-Annotations2audio10M<n<100M1 likes318 downloads1y agoHugging Face12TTS-AGI /emolia-hq Emolia-HQ Emolia-HQ is a high-quality, speaker-paired subset of the LAION Emolia dataset. Each sample includes a target utterance and a reference utterance from the same speaker, enabling speaker-conditioned tasks such as voice conversion, expressive TTS, and speaker-aware emotion recognition. Source Derived from laion/Emolia by: Quality filtering: Only samples with dnsmos >= 3.0 are retained. Speaker pairing: Each target sample is matched with a reference audio from the… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emolia-hq.audioaudio-classification10M<n<100M3 likes155 downloads7mo agoHugging Face13AffectDF /AffectDF_EmotionSDD AffectDF: Emotionally Expressive Speech Deepfake Benchmark Overview AffectDF is a large-scale benchmark for speech deepfake detection under emotionally expressive spoofing conditions. The dataset is designed to evaluate whether current speech deepfake detection (SDD) systems can generalize beyond conventional neutral-speech benchmarks to modern emotional and expressive speech attacks. AffectDF contains approximately 260 hours of audio generated using 21 spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AffectDF/AffectDF_EmotionSDD.audioaudio-classification100K<n<1M0 likes140 downloads4mo agoHugging Face14TTS-AGI /Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave Emotion and Voice Attribute Reference Snippets - DACVAE and Wave Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio. Overview Total samples: 606,178 Filtered out: 363,331 (samples with speech_quality < 1.8) Total tar files: 328 Total size: 1.54 TB Audio format: WAV, 48kHz, PCM 16-bit mono Latents: DAC-VAE float16 [T, 128] at 25 frames/sec Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.audiotext-to-speech100K<n<1M0 likes86 downloads6mo agoHugging Face15laion /emolia-3k-speaker-clusters Emolia 3K Speaker Clusters A curated set of 3,000 diverse speaker clusters derived from the TTS-AGI/emolia-hq dataset, with up to 20 representative audio samples per cluster. Overview The original emolia-hq dataset contains hundreds of thousands of speech samples with 128-dimensional WavLM speaker timbre embeddings. These were first clustered into 10,000 centroids, then intelligently pruned to 3,000 using density-aware farthest-point sampling to ensure: Outlier… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-3k-speaker-clusters.audioaudio-classification10K<n<100K0 likes65 downloads6mo agoHugging Face16laion /emolia-balanced-5M-subset emolia-balanced-5M-subset A balanced ~5.26M-sample subset of laion/Emolia (80.5M speech samples), packaged as WebDataset-compatible tar shards for direct use in training pipelines. How this subset was filtered Samples were selected if they met either of two criteria: 1. Emotion thresholds Each sample carries 40 emotion annotation scores (from the Emonet taxonomy) in its metadata. A sample qualifies for an emotion bucket if its score for that emotion meets or… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-balanced-5M-subset.audio1M<n<10M1 likes43 downloads5mo agoHugging Face17emoji-tts /emoji-tts-22k Emoji-TTS 22K Training Data Emoji-TTS 22K is the training corpus used to build Emoji-TTS, an emoji-conditioned expressive text-to-speech model. Each example pairs an English transcript, an emoji control label, and a synthetic WAV utterance spoken with the fixed Kore voice. The corpus contains 21,940 utterances: 19,945 emoji-conditioned samples and 1,995 neutral/no-emoji samples. The control inventory covers ten emoji labels plus the neutral <none> label. How It Was Built… See the full description on the dataset page: https://huggingface.co/datasets/emoji-tts/emoji-tts-22k.audiotext-to-speech10K<n<100K0 likes36 downloads5mo agoHugging Face18TTS-AGI /voice-emo-cloning-dataset Emotion-Cloning TTS Training Dataset Location /home/deployer/laion/echo-tts-training-main/emotion_eval/dataset_output/ Overview This dataset contains ~22,518 training triplets for fine-tuning a zero-shot voice+emotion cloning TTS model. Each sample provides everything needed to train a model that can clone both a speaker's voice identity AND their emotional delivery from separate reference audio clips. The data is stored as WebDataset .tar shards, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-emo-cloning-dataset.audio10K<n<100K0 likes29 downloads6mo agoHugging Face19laion /en_and_de_reference_voice_files_for_emotion_cloningaudio100K<n<1M2 likes13 downloads1y agoHugging Face20krishnakalyan3 /emo_subset_webdsaudio10K<n<100K0 likes11 downloads2y agoHugging Face21TTS-AGI /Emotion-Voice-Attribute-Reference-Snippets-DACVAE Emotion and Voice Attribute Reference Snippets - DACVAE and Wave Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio. Overview Total samples: 606,178 Filtered out: 363,331 (samples with speech_quality < 1.8) Total tar files: 328 Total size: ~98 GB (latents-only, no WAV) Audio format: WAV, 48kHz, PCM 16-bit mono Latents: DAC-VAE float16 [T, 128] at 25 frames/sec Dimensions: 57 (40… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE.texttext-to-speech100K<n<1M0 likes10 downloads6mo agoHugging Face22krishnakalyan3 /emo_speech_sampleaudio1K<n<10K1 likes6 downloads2y agoHugging Face23jspaulsen /emotional-tts-wikiaudio10K<n<100K0 likes6 downloads6mo agoHugging Face24cahya /emo-smallaudio1K<n<10K0 likes1 downloads2y agoHugging Face25mrfakename /EmoAct-SFT-Datagatedaudio100K<n<1M0 likes1 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.