CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /Emolia Dataset Card for Emolia Dataset Description This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?". The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.audio10M<n<100M15 likes2.1k downloads10mo agoHugging Face02krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face03krishnakalyan3 /emo_parleraudio1M<n<10M2 likes1.4k downloads2y agoHugging Face04krishnakalyan3 /emo_webdsaudio10K<n<100K5 likes1.3k downloads2y agoHugging Face05laion /Emilia-with-Emotion-Annotations4audio10M<n<100M1 likes895 downloads1y agoHugging Face06akuzdeuov /qwen3-tts-multilingual-emotional-speechaudio1M<n<10M0 likes748 downloads14d agoHugging Face07laion /Emilia-with-Emotion-Annotations5audio10M<n<100M3 likes735 downloads1y agoHugging Face08krishnakalyan3 /emo_speech_filtered_v12 second filtered emotional speech in webdataset format https://huggingface.co/datasets/EQ4You/Emotional_Speech audio10K<n<100K0 likes634 downloads2y agoHugging Face09BAAI /Emotiontalkgated EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations Introduction EmotionTalk is an interactive Chinese multimodal emotion dataset with rich annotations. This dataset provides multimodal information from 19 actors participating in dyadic conversation settings, incorporating acoustic, visual, and textual modalities. It includes 23.6 hours of speech (19,250 utterances), annotations for 7 utterance-level emotion categories (happy… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Emotiontalk.text100K<n<1M34 likes442 downloads1y agoHugging Face10laion /Emilia-with-Emotion-Annotations3audio10M<n<100M1 likes383 downloads1y agoHugging Face11VoiceNet /emolia emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired) This is the emolia-balanced-5M-subset corpus repackaged for high-quality audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz (PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json samples. The JSON sidecar carries the full annotation stack: Original metadata (id, text, duration, speaker, language, dnsmos). A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.audioaudio-classification1M<n<10M1 likes381 downloads5mo agoHugging Face12laion /Emilia-with-Emotion-Annotations2audio10M<n<100M1 likes318 downloads1y agoHugging Face13laion /emonet-face-big EmoNet-Face: A Fine-Grained, Expert-Annotated Benchmark for Facial Emotion Recognition Dataset Summary EmoNet-Face is a comprehensive benchmark suite designed to address critical gaps in facial emotion recognition (FER). Current benchmarks often have a narrow emotional spectrum, lack demographic diversity, and use uncontrolled imagery. EmoNet-Face provides a robust foundation for developing and evaluating AI systems with a deeper, more nuanced understanding of human… See the full description on the dataset page: https://huggingface.co/datasets/laion/emonet-face-big.image100K<n<1M10 likes303 downloads11mo agoHugging Face14printblue /EmoArt-130k EmoArt: A Large-Scale Emotion-Annotated Artistic Dataset Overview EmoArt is a comprehensive, large-scale emotion-annotated artistic dataset containing 132,664 high-resolution artworks spanning 56 painting styles across 7 thematic categories. This dataset bridges the gap between visual art and emotional computing, enabling groundbreaking research in emotion-aware AI systems. Key Statistics 📊 132,664 artworks with rich emotional annotations 🎨 56 distinct… See the full description on the dataset page: https://huggingface.co/datasets/printblue/EmoArt-130k.image100K<n<1M9 likes296 downloads1y agoHugging Face15TTS-AGI /emolia-hq Emolia-HQ Emolia-HQ is a high-quality, speaker-paired subset of the LAION Emolia dataset. Each sample includes a target utterance and a reference utterance from the same speaker, enabling speaker-conditioned tasks such as voice conversion, expressive TTS, and speaker-aware emotion recognition. Source Derived from laion/Emolia by: Quality filtering: Only samples with dnsmos >= 3.0 are retained. Speaker pairing: Each target sample is matched with a reference audio from the… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emolia-hq.audioaudio-classification10M<n<100M3 likes155 downloads7mo agoHugging Face16AffectDF /AffectDF_EmotionSDD AffectDF: Emotionally Expressive Speech Deepfake Benchmark Overview AffectDF is a large-scale benchmark for speech deepfake detection under emotionally expressive spoofing conditions. The dataset is designed to evaluate whether current speech deepfake detection (SDD) systems can generalize beyond conventional neutral-speech benchmarks to modern emotional and expressive speech attacks. AffectDF contains approximately 260 hours of audio generated using 21 spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AffectDF/AffectDF_EmotionSDD.audioaudio-classification100K<n<1M0 likes140 downloads4mo agoHugging Face17TTS-AGI /Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave Emotion and Voice Attribute Reference Snippets - DACVAE and Wave Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio. Overview Total samples: 606,178 Filtered out: 363,331 (samples with speech_quality < 1.8) Total tar files: 328 Total size: 1.54 TB Audio format: WAV, 48kHz, PCM 16-bit mono Latents: DAC-VAE float16 [T, 128] at 25 frames/sec Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.audiotext-to-speech100K<n<1M0 likes86 downloads6mo agoHugging Face18TTS-AGI /emolia-3k-speaker-clusters-DACVAE Emolia 3K Speaker Clusters A curated set of 3,000 diverse speaker clusters derived from the TTS-AGI/emolia-hq dataset, with up to 20 representative audio samples per cluster. Overview The original emolia-hq dataset contains hundreds of thousands of speech samples with 128-dimensional WavLM speaker timbre embeddings. These were first clustered into 10,000 centroids, then intelligently pruned to 3,000 using density-aware farthest-point sampling to ensure: Outlier… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emolia-3k-speaker-clusters-DACVAE.textaudio-classification10K<n<100K0 likes85 downloads6mo agoHugging Face19mengtingwei /emotion_bias Emotion Bias in Synthetic Face Generation Description This dataset accompanies the paper "Happy Young Women, Grumpy Old Men? Emotion Prompts as Demographic Selectors in AI Image Generation". It contains 56,000 synthetic face images generated by eight state-of-the-art text-to-image (T2I) models across seven emotion prompt conditions, along with demographic attribute annotations (gender, race, age) and perceived attractiveness labels for each image. The dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/mengtingwei/emotion_bias.imageimage-classification10K<n<100K1 likes77 downloads5mo agoHugging Face20laion /emolia-3k-speaker-clusters Emolia 3K Speaker Clusters A curated set of 3,000 diverse speaker clusters derived from the TTS-AGI/emolia-hq dataset, with up to 20 representative audio samples per cluster. Overview The original emolia-hq dataset contains hundreds of thousands of speech samples with 128-dimensional WavLM speaker timbre embeddings. These were first clustered into 10,000 centroids, then intelligently pruned to 3,000 using density-aware farthest-point sampling to ensure: Outlier… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-3k-speaker-clusters.audioaudio-classification10K<n<100K0 likes65 downloads6mo agoHugging Face21laion /emolia-balanced-5M-subset emolia-balanced-5M-subset A balanced ~5.26M-sample subset of laion/Emolia (80.5M speech samples), packaged as WebDataset-compatible tar shards for direct use in training pipelines. How this subset was filtered Samples were selected if they met either of two criteria: 1. Emotion thresholds Each sample carries 40 emotion annotation scores (from the Emonet taxonomy) in its metadata. A sample qualifies for an emotion bucket if its score for that emotion meets or… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-balanced-5M-subset.audio1M<n<10M1 likes43 downloads5mo agoHugging Face22emoji-tts /emoji-tts-22k Emoji-TTS 22K Training Data Emoji-TTS 22K is the training corpus used to build Emoji-TTS, an emoji-conditioned expressive text-to-speech model. Each example pairs an English transcript, an emoji control label, and a synthetic WAV utterance spoken with the fixed Kore voice. The corpus contains 21,940 utterances: 19,945 emoji-conditioned samples and 1,995 neutral/no-emoji samples. The control inventory covers ten emoji labels plus the neutral <none> label. How It Was Built… See the full description on the dataset page: https://huggingface.co/datasets/emoji-tts/emoji-tts-22k.audiotext-to-speech10K<n<100K0 likes36 downloads5mo agoHugging Face23TTS-AGI /voice-emo-cloning-dataset Emotion-Cloning TTS Training Dataset Location /home/deployer/laion/echo-tts-training-main/emotion_eval/dataset_output/ Overview This dataset contains ~22,518 training triplets for fine-tuning a zero-shot voice+emotion cloning TTS model. Each sample provides everything needed to train a model that can clone both a speaker's voice identity AND their emotional delivery from separate reference audio clips. The data is stored as WebDataset .tar shards, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-emo-cloning-dataset.audio10K<n<100K0 likes29 downloads6mo agoHugging Face24TTS-AGI /enhanced-emo-snippets-balanced-DACVAE Enhanced Emotion Snippets — Balanced DACVAE A balanced, emotion-bucketed subset of TTS-AGI/enhanced-audiosnippets-DACVAE, organized by Empathic Insight Voice+ emotion and voice attribute categories. Overview This dataset provides up to 100 samples per magnitude bucket for each of the 40 emotion categories and 15 voice attribute dimensions scored by Empathic Insight Voice+. Selection Criteria Emotion Categories (40 dimensions) For each emotion (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/enhanced-emo-snippets-balanced-DACVAE.textaudio-classification10K<n<100K0 likes26 downloads6mo agoHugging Face25chitradrishti /Emoticimage10K<n<100K1 likes24 downloads3y agoHugging Face26TTS-AGI /emotion-attribute-conditioning-dacvae Echo TTS - Emotion & Attribute Conditioning Dataset (DAC-VAE Latents) Pre-bucketed speech dataset with DAC-VAE latent representations organized by 40 emotion categories and 13 vocal/audio attributes. Built for conditioning fine-tuning of Echo TTS and similar DiT-based TTS models. Overview Total emotion samples: 163,271 (across 40 emotions, 10K cap per emotion) Total attribute samples: ~785K (across 13 attributes x 7 buckets, 10K cap per bucket) Format: WebDataset .tar… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emotion-attribute-conditioning-dacvae.text1M<n<10M0 likes16 downloads7mo agoHugging Face27laion /en_and_de_reference_voice_files_for_emotion_cloningaudio100K<n<1M2 likes13 downloads1y agoHugging Face28smallbraineng /voxbox-vb-wds-emotion-fulltext100K<n<1M0 likes13 downloads1y agoHugging Face29krishnakalyan3 /emo_subset_webdsaudio10K<n<100K0 likes11 downloads2y agoHugging Face30TTS-AGI /Emotion-Voice-Attribute-Reference-Snippets-DACVAE Emotion and Voice Attribute Reference Snippets - DACVAE and Wave Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio. Overview Total samples: 606,178 Filtered out: 363,331 (samples with speech_quality < 1.8) Total tar files: 328 Total size: ~98 GB (latents-only, no WAV) Audio format: WAV, 48kHz, PCM 16-bit mono Latents: DAC-VAE float16 [T, 128] at 25 frames/sec Dimensions: 57 (40… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE.texttext-to-speech100K<n<1M0 likes10 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.