datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-with-Emotion-Annotations4Emilia-with-Emotion-Annotations5qwen3-tts-multilingual-emotional-speechEmotiontalk
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
Introduction
EmotionTalk is an interactive Chinese multimodal emotion dataset with rich annotations. This dataset provides multimodal information from 19 actors participating in dyadic conversation settings, incorporating acoustic, visual, and textual modalities. It includes 23.6 hours of speech (19,250 utterances), annotations for 7 utterance-level emotion categories (happy… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Emotiontalk.Emilia-with-Emotion-Annotations3Emilia-with-Emotion-Annotations2AffectDF_EmotionSDD
AffectDF: Emotionally Expressive Speech Deepfake Benchmark
Overview
AffectDF is a large-scale benchmark for speech deepfake detection under emotionally expressive spoofing conditions. The dataset is designed to evaluate whether current speech deepfake detection (SDD) systems can generalize beyond conventional neutral-speech benchmarks to modern emotional and expressive speech attacks.
AffectDF contains approximately 260 hours of audio generated using 21 spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AffectDF/AffectDF_EmotionSDD.Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: 1.54 TB
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.emotion_bias
Emotion Bias in Synthetic Face Generation
Description
This dataset accompanies the paper "Happy Young Women, Grumpy Old Men? Emotion Prompts as Demographic Selectors in AI Image Generation".
It contains 56,000 synthetic face images generated by eight state-of-the-art text-to-image (T2I) models across seven emotion prompt conditions, along with demographic attribute annotations (gender, race, age) and perceived attractiveness labels for each image.
The dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/mengtingwei/emotion_bias.emotion-attribute-conditioning-dacvae
Echo TTS - Emotion & Attribute Conditioning Dataset (DAC-VAE Latents)
Pre-bucketed speech dataset with DAC-VAE latent representations organized by 40 emotion categories and 13 vocal/audio attributes. Built for conditioning fine-tuning of Echo TTS and similar DiT-based TTS models.
Overview
Total emotion samples: 163,271 (across 40 emotions, 10K cap per emotion)
Total attribute samples: ~785K (across 13 attributes x 7 buckets, 10K cap per bucket)
Format: WebDataset .tar… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emotion-attribute-conditioning-dacvae.voxbox-vb-wds-emotion-fullen_and_de_reference_voice_files_for_emotion_cloningvoxbox-vb-wds-emotion-2Emotion-Voice-Attribute-Reference-Snippets-DACVAE
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: ~98 GB (latents-only, no WAV)
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE.latent-toronto-emotionThe Toronto Emotional Speech Set (TESS) is a dataset consisting of emotionally charged speech recordings, designed for emotion recognition tasks. This repository provides precomputed audio embeddings extracted using the Music2Latent model. These embeddings are derived from the TESS dataset, available at TESS dataset on Kaggle, enabling quick and efficient use for tasks like speech emotion recognition. The embeddings can be directly used for classification tasks, without the need for raw audio… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/latent-toronto-emotion.emotional-tts-wikivoxbox-vb-wds-emotionvoxbox-vb-wds-emotion-full-2
