datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emova-alignment-7m
EMOVA-Alignment-7M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment.
This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data.
This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.emova-sft-4m
EMOVA-SFT-4M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.emolia
emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired)
This is the emolia-balanced-5M-subset corpus repackaged for high-quality
audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz
(PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json
samples.
The JSON sidecar carries the full annotation stack:
Original metadata (id, text, duration, speaker, language, dnsmos).
A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.emova-sft-speech-231k
EMOVA-SFT-Speech-231K
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-SFT-Speech-231K is a comprehensive dataset curated for omni-modal instruction tuning and emotional spoken dialogue. This dataset is created by converting existing text and visual instruction datasets via Text-to-Speech (TTS) tools. EMOVA-SFT-Speech-231K is part of EMOVA-Datasets collection and is used in… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-speech-231k.emova-asr-tts-eval
EMOVA-ASR-TTS-Eval
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-ASR-TTS-Eval is a dataset designed for evaluating the ASR and TTS performance of Omni-modal LLMs. It is derived from the test-clean set of the LibriSpeech dataset. This dataset is part of the EMOVA-Datasets collection. We extract the speech units using the EMOVA Speech Tokenizer.
Structure
This… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-asr-tts-eval.arabic-multidialect-emotional-speech-demo
DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech
A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request.
Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.egyption-with-emotion-dataset
Egption Text-Audio Dataset With Emotions and Diarization
Creating datasets for TTS and ASR models with emotions and Diarization
In case you want to focus only one speaker , you can fiter based on speaker_role
Source Code
if you want to collect more data from youtube, you can check this link
🙏 Acknowledgements
This project makes use of the forced alignment model and Cohere ASR model provided by:
MahmoudAshraf/mms-300m-1130-forced-aligner
Cohere ASR
Hubert… See the full description on the dataset page: https://huggingface.co/datasets/OmarAhmedSobhy/egyption-with-emotion-dataset.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.TTS-emotion-voice
TTS Emotion Voice: Unconfident Interview Speech
以台灣華語面試回答文字合成的不自信語音資料集。這是研究用的合成資料,不是真人錄音;
所有 unconfident=1 標籤來自生成 prompt,尚未經完整的人類情緒標註驗證。
兩個設定
config
筆數
取樣率
時長範圍
總時長
用途
max12
471
24 kHz
6.267–12.000 秒
4,636.990 秒
保留較完整語意與韻律的長片段對照組
nnime_match_v2
471
16 kHz
0.256–12.000 秒
1,188.188 秒
匹配 NNIME Train Unconfident 時長分布的硬切消融組
兩個 config 使用相同 471 個來源 parent、相同 U25/U50/U75 成員與 prompt 配置;
U25 包含於 U50,U50 包含於 U75。metadata.csv 的 in_u25、in_u50、in_u75… See the full description on the dataset page: https://huggingface.co/datasets/guan-chen/TTS-emotion-voice.emo-com
Emo-con
Emo-con is a speech dataset of real emotional conversations between people actively supporting each other. Unlike most emotion datasets, which rely on acted, pseudo-acted, or scripted speech, Emo-con captures genuine emotional expression in the context of mutual support. It reflects the way people actually talk to a close friend when sharing hardships, daily experiences, and jokes.
We operate support and community groups with licensed professionals, so we can assume all… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/emo-com.Arabic-Emotional-Audio-Dataset-Baved
BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging)
A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits.
Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.emo_tts
Tatar Dubbed Speech
Sentence-level speech segments in Tatar, cut from Tatar-language dubs and aligned to their
subtitles with CTC forced alignment.
7,257 sentences / 4:39:06 drawn from 11.85 h of source audio — dialogue is sparse in this
material, so roughly 40% of the runtime is speech.
Source
source prefix
Rows
Duration
Median similarity
Берсерк, 25 episodes
Берсерк - N серия
5,855
4:03:34
0.941
Мистер һәм миссис Смит
Мистер һәм миссис Смит
1,402
0:35:32
0.909… See the full description on the dataset page: https://huggingface.co/datasets/gaydmi/emo_tts.emova-sft-speech-eval
EMOVA-SFT-Speech-Eval
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-SFT-Speech-Eval is an evaluation dataset curated for omni-modal instruction tuning and emotional spoken dialogue. This dataset is created by converting existing text and visual instruction datasets via Text-to-Speech (TTS) tools. EMOVA-SFT-Speech-Eval is part of EMOVA-Datasets collection, and the training… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-speech-eval.eng-sports-radio-psst-iu-emotion-splits
English Sports Radio Non-Neutral Emotion IU Splits
Public non-neutral subset of NathanRoll/eng-sports-radio-psst-iu.
Each row is one intonation unit with exactly three columns:
audio: embedded 16 kHz mono audio for the IU
text: a leading emotion special token followed by the Parakeet transcript
accent: broadcast-location proxy accent label
Neutral examples were removed. The remaining rows are split by emotion:
joy: 247 rows, 0.287 audio hours
surprise: 153 rows, 0.188 audio… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/eng-sports-radio-psst-iu-emotion-splits.indian-tts-emotion-60min
indian-tts-emotion-60min
A small, carefully curated text-to-speech dataset: ~68 minutes of clean,
single-speaker-per-clip audio in Indian English (en-IN) and Hindi (hi-IN), sourced
from YouTube, with accurate transcriptions and per-clip emotion/style tags.
Built as a data-quality exercise: clips were filtered conservatively and a sample was
verified by listening rather than shipped straight from an automated pipeline.
Summary
Language
Clips
Duration (min)… See the full description on the dataset page: https://huggingface.co/datasets/sarthwa8/indian-tts-emotion-60min.audio-emotion-detection-dataset
Audio Emotion Detection Dataset
Github: Audio Emotion Detection Dataset
Connect with me : Linkedin
Speech clips in English and Hindi annotated with emotion labels and ASR transcripts.
Audio is sourced from public YouTube videos and trimmed to approximately 60 seconds per clip.
Noise reduction is applied via noisereduce and silero-vad.
Emotions (5 classes)
Label
Description
angry
Aggressive, confrontational speech
calm… See the full description on the dataset page: https://huggingface.co/datasets/kotangalechaitali9007/audio-emotion-detection-dataset.raw-emocean
raw-emocean
Large-scale English speech dataset for text-to-speech (TTS) model training. Designed for autoregressive TTS architectures (TADA, CSM, VALL-E style models).
Dataset Summary
Metric
Value
Parquet shards
7
Segment duration
3–8 seconds
Sample rate
24,000 Hz (mono)
ASR engine
NVIDIA Parakeet TDT 0.6B v3
Format
Parquet with embedded audio
Dataset Schema
Column
Type
Description
audio
Audio
Waveform array + sampling… See the full description on the dataset page: https://huggingface.co/datasets/somu9/raw-emocean.
