datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.emotion-preserving-s2st-benchmark
Emotion-Preserving Speech-to-Speech Translation Benchmark
Overview
First benchmark for evaluating emotion preservation in speech-to-speech translation systems.
Pipeline: English speech → Emotion Detection → Translation (EN→HI) → Emotion-conditioned TTS.
Key Results (1440 samples, RAVDESS dataset)
Metric
Score
Emotion Detection Accuracy
36.3% (4-class on 8-class data)
Emotion Preservation Rate
43.5%
Preservation Gain over Flat TTS
+7.2%
F0… See the full description on the dataset page: https://huggingface.co/datasets/primal-sage/emotion-preserving-s2st-benchmark.voxtral-emotion-temporal
VoxTral Emotion Temporal Dataset
Dataset for training emotion transition detection in speech. ~500 clips with frame-level emotion annotations at 20ms resolution.
Overview
Property
Value
Clips
~500
Sample Rate
16000 Hz
Frame Resolution
20ms
Emotions
6 (neutral, happy, angry, sad, surprise, fear)
Clip Distribution
Single Emotion (40%): One emotion throughout
Transitions (40%): Emotion changes mid-sentence (2-3 segments)
No Emotion (20%):… See the full description on the dataset page: https://huggingface.co/datasets/MrlolDev/voxtral-emotion-temporal.
