datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t9p3c8m1-axr4e6vocal-affect-bench
VocalAffectBench
VocalAffectBench is a test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio.
Paper: VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models
The benchmark targets the expressed emotion — what the speaker conveys through vocal tone, prosody, pace, intensity, and pauses — not inferred internal state.
Contents
280 human-recorded English WAV clips, totalling 2.32 hours.
7… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/vocal-affect-bench.h4p7t3x2-jn6b9_tran
affectexpect/h4p7t3x2-jn6b9_tran
This dataset contains transcribed audio files organized in folders for scalability.
Dataset Structure
The dataset is organized with:
Audio files: Stored in audio_XXXXX/ folders (5000 files per folder)
Metadata: Stored in data_XXXXX/ folders as parquet files
This organization follows Hugging Face best practices for datasets with millions of files.
Statistics
Total files: 926
Total batches: 2427
Audio folders: 3
Files per… See the full description on the dataset page: https://huggingface.co/datasets/affectexpect/h4p7t3x2-jn6b9_tran.t9p3c8m1-axr4e6_sepAffectDF_EmotionSDD
AffectDF: Emotionally Expressive Speech Deepfake Benchmark
Overview
AffectDF is a large-scale benchmark for speech deepfake detection under emotionally expressive spoofing conditions. The dataset is designed to evaluate whether current speech deepfake detection (SDD) systems can generalize beyond conventional neutral-speech benchmarks to modern emotional and expressive speech attacks.
AffectDF contains approximately 260 hours of audio generated using 21 spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AffectDF/AffectDF_EmotionSDD.m4r9e1x8-cd7h2n4x7d2q9-hf1m8t3_sept9p3c8m1-axr4e6_tran
affectexpect/t9p3c8m1-axr4e6_tran
This dataset contains transcribed audio files organized in folders for scalability.
Dataset Structure
The dataset is organized with:
Audio files: Stored in audio_XXXXX/ folders (5000 files per folder)
Metadata: Stored in data_XXXXX/ folders as parquet files
This organization follows Hugging Face best practices for datasets with millions of files.
Statistics
Total files: 8,901
Total batches: 5183
Audio folders: 6
Files per… See the full description on the dataset page: https://huggingface.co/datasets/affectexpect/t9p3c8m1-axr4e6_tran.h4p7t3x2-jn6b9_sepn4x7d2q9-hf1m8t3_tran
affectexpect/n4x7d2q9-hf1m8t3_tran
This dataset contains transcribed audio files organized in folders for scalability.
Dataset Structure
The dataset is organized with:
Audio files: Stored in audio_XXXXX/ folders (5000 files per folder)
Metadata: Stored in data_XXXXX/ folders as parquet files
This organization follows Hugging Face best practices for datasets with millions of files.
Statistics
Total files: 922
Total batches: 11336
Audio folders: 10… See the full description on the dataset page: https://huggingface.co/datasets/affectexpect/n4x7d2q9-hf1m8t3_tran.AffectHuman-43K
AffectHuman-43K
AffectHuman-43K is an emotion-aligned multimodal benchmark for controlled human affect generation and evaluation.
The benchmark contains 42,469 usable samples with complete image, reference-image, audio, and text coverage. Identity is specified through a visual reference image, while text, audio, and emotion labels provide affective control signals. This design separates identity preservation from affective control, enabling evaluation of whether a model can preserve… See the full description on the dataset page: https://huggingface.co/datasets/iamjamuna/AffectHuman-43K.h4p7t3x2-jn6b9_embeddedn4x7d2q9-hf1m8t3_embeddedt9p3c8m1-axr4e6_embeddedb3h9r7k2-mp1x6c8x4n9r1-mj7t2s2m4h9t7-lc8p0h4p7t3x2-jn6b9q2k8f3n1-lr9t0m
