CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01VoiceNet /emolia-thinking Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.audioaudio-classification100K<n<1M0 likes5.1k downloads3mo agoHugging Face02langswap /dialogs-ru-emotional-conversations Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational Russian speech, designed for dialog-oriented and emotional text-to-speech. Unlike existing Russian corpora — mostly single-speaker read speech or large but low-quality web-mined audio — Dialogs was recorded by professional theatre actors performing scripted dialogs face-to-face, capturing natural turn-taking, timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.audiotext-to-speechn<1K18 likes2.4k downloads2mo agoHugging Face03VoiceNet /emolia emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired) This is the emolia-balanced-5M-subset corpus repackaged for high-quality audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz (PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json samples. The JSON sidecar carries the full annotation stack: Original metadata (id, text, duration, speaker, language, dnsmos). A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.audioaudio-classification1M<n<10M1 likes381 downloads5mo agoHugging Face04UniDataPro /speech-emotion-recognition Speech Emotion Recognition Dataset comprises 30,000+ audio recordings featuring 4 distinct emotions: euphoria, joy, sadness, and surprise. This extensive collection is designed for research in emotion recognition, focusing on the nuances of emotional speech and the subtleties of speech signals as individuals vocally express their feelings. By utilizing this dataset, researchers and developers can enhance their understanding of sentiment analysis and improve automatic speech… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/speech-emotion-recognition.audioautomatic-speech-recognitionn<1K6 likes169 downloads1mo agoHugging Face05datahiveai /arabic-multidialect-emotional-speech-demo DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request. Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.audioautomatic-speech-recognitionn<1K3 likes122 downloads5mo agoHugging Face06OmarAhmedSobhy /egyption-with-emotion-dataset Egption Text-Audio Dataset With Emotions and Diarization Creating datasets for TTS and ASR models with emotions and Diarization In case you want to focus only one speaker , you can fiter based on speaker_role Source Code if you want to collect more data from youtube, you can check this link 🙏 Acknowledgements This project makes use of the forced alignment model and Cohere ASR model provided by: MahmoudAshraf/mms-300m-1130-forced-aligner Cohere ASR Hubert… See the full description on the dataset page: https://huggingface.co/datasets/OmarAhmedSobhy/egyption-with-emotion-dataset.audioautomatic-speech-recognition1K<n<10K4 likes105 downloads5mo agoHugging Face07RapidOrc121 /audio-emotion-detection-dataset Audio Emotion Detection Dataset Github: Audio Emotion Detection Dataset Connect with me : Linkedin Speech clips in English and Hindi annotated with emotion labels and ASR transcripts. Audio is sourced from public YouTube videos and trimmed to approximately 60 seconds per clip. Noise reduction is applied via noisereduce and silero-vad. Emotions (5 classes) Label Description angry Aggressive, confrontational speech calm… See the full description on the dataset page: https://huggingface.co/datasets/RapidOrc121/audio-emotion-detection-dataset.audioaudio-classification1 likes103 downloads4mo agoHugging Face08FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes86 downloads5mo agoHugging Face09guan-chen /TTS-emotion-voice TTS Emotion Voice: Unconfident Interview Speech 以台灣華語面試回答文字合成的不自信語音資料集。這是研究用的合成資料,不是真人錄音; 所有 unconfident=1 標籤來自生成 prompt,尚未經完整的人類情緒標註驗證。 兩個設定 config 筆數 取樣率 時長範圍 總時長 用途 max12 471 24 kHz 6.267–12.000 秒 4,636.990 秒 保留較完整語意與韻律的長片段對照組 nnime_match_v2 471 16 kHz 0.256–12.000 秒 1,188.188 秒 匹配 NNIME Train Unconfident 時長分布的硬切消融組 兩個 config 使用相同 471 個來源 parent、相同 U25/U50/U75 成員與 prompt 配置; U25 包含於 U50,U50 包含於 U75。metadata.csv 的 in_u25、in_u50、in_u75… See the full description on the dataset page: https://huggingface.co/datasets/guan-chen/TTS-emotion-voice.audioaudio-classificationn<1K0 likes62 downloads13d agoHugging Face10liva-ai /emo-comgated Emo-con Emo-con is a speech dataset of real emotional conversations between people actively supporting each other. Unlike most emotion datasets, which rely on acted, pseudo-acted, or scripted speech, Emo-con captures genuine emotional expression in the context of mutual support. It reflects the way people actually talk to a close friend when sharing hardships, daily experiences, and jokes. We operate support and community groups with licensed professionals, so we can assume all… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/emo-com.audioaudio-classificationn<1K1 likes55 downloads4mo agoHugging Face11FatimahEmadEldin /Arabic-Emotional-Audio-Dataset-Baved BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging) A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits. Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.audioaudio-classification1K<n<10K0 likes48 downloads5mo agoHugging Face12gaydmi /emo_ttsgated Tatar Dubbed Speech Sentence-level speech segments in Tatar, cut from Tatar-language dubs and aligned to their subtitles with CTC forced alignment. 7,257 sentences / 4:39:06 drawn from 11.85 h of source audio — dialogue is sparse in this material, so roughly 40% of the runtime is speech. Source source prefix Rows Duration Median similarity Берсерк, 25 episodes Берсерк - N серия 5,855 4:03:34 0.941 Мистер һәм миссис Смит Мистер һәм миссис Смит 1,402 0:35:32 0.909… See the full description on the dataset page: https://huggingface.co/datasets/gaydmi/emo_tts.audiotext-to-speech1K<n<10K0 likes40 downloads5d agoHugging Face13NathanRoll /eng-sports-radio-psst-iu-emotion-splits English Sports Radio Non-Neutral Emotion IU Splits Public non-neutral subset of NathanRoll/eng-sports-radio-psst-iu. Each row is one intonation unit with exactly three columns: audio: embedded 16 kHz mono audio for the IU text: a leading emotion special token followed by the Parakeet transcript accent: broadcast-location proxy accent label Neutral examples were removed. The remaining rows are split by emotion: joy: 247 rows, 0.287 audio hours surprise: 153 rows, 0.188 audio… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/eng-sports-radio-psst-iu-emotion-splits.audioautomatic-speech-recognitionn<1K0 likes33 downloads4mo agoHugging Face14TTS-AGI /moss-emolia-elise-hq-captioned MOSS · Emolia + Elise + Inline-Bursts — HQ, captioned A high-quality, richly captioned slice of the MOSS-local voice-acting corpus: expressive speech clips scored by a panel of acoustic detectors, filtered to the top by a composite reward, and captioned in the voice-acting format (a "how the voice sounds / how to perform it" description plus the script with inline vocal-burst tags). Audio is shipped both as flac (WebDataset tars) and as pre-computed MOSS-Audio-Tokenizer codes… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-emolia-elise-hq-captioned.audiotext-to-speech1K<n<10K0 likes32 downloads2mo agoHugging Face15pan82 /bedvibe-emotional-speechBedVibe Emotional Speech Dataset Studio-quality emotional speech dataset • Preview samples available on Hugging Face • Full commercial dataset available via BedVibe Studio • Languages currently available: English, Greek • Additional languages can be recorded on request • 6 emotions • 48 kHz / 32-bit float audio • Professionally recorded in studio conditions Studio-recorded emotional speech datasets designed for training text-to-speech systems. Currently available on our website: more than 108… See the full description on the dataset page: https://huggingface.co/datasets/pan82/bedvibe-emotional-speech.audiotext-to-speechn<1K0 likes20 downloads7mo agoHugging Face16sarthwa8 /indian-tts-emotion-60min indian-tts-emotion-60min A small, carefully curated text-to-speech dataset: ~68 minutes of clean, single-speaker-per-clip audio in Indian English (en-IN) and Hindi (hi-IN), sourced from YouTube, with accurate transcriptions and per-clip emotion/style tags. Built as a data-quality exercise: clips were filtered conservatively and a sample was verified by listening rather than shipped straight from an automated pipeline. Summary Language Clips Duration (min)… See the full description on the dataset page: https://huggingface.co/datasets/sarthwa8/indian-tts-emotion-60min.audiotext-to-speechn<1K1 likes20 downloads3mo agoHugging Face17kotangalechaitali9007 /audio-emotion-detection-dataset Audio Emotion Detection Dataset Github: Audio Emotion Detection Dataset Connect with me : Linkedin Speech clips in English and Hindi annotated with emotion labels and ASR transcripts. Audio is sourced from public YouTube videos and trimmed to approximately 60 seconds per clip. Noise reduction is applied via noisereduce and silero-vad. Emotions (5 classes) Label Description angry Aggressive, confrontational speech calm… See the full description on the dataset page: https://huggingface.co/datasets/kotangalechaitali9007/audio-emotion-detection-dataset.audioaudio-classificationn<1K0 likes16 downloads1mo agoHugging Face18somu9 /raw-emoceangated raw-emocean Large-scale English speech dataset for text-to-speech (TTS) model training. Designed for autoregressive TTS architectures (TADA, CSM, VALL-E style models). Dataset Summary Metric Value Parquet shards 7 Segment duration 3–8 seconds Sample rate 24,000 Hz (mono) ASR engine NVIDIA Parakeet TDT 0.6B v3 Format Parquet with embedded audio Dataset Schema Column Type Description audio Audio Waveform array + sampling… See the full description on the dataset page: https://huggingface.co/datasets/somu9/raw-emocean.audiotext-to-speech10K<n<100K1 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.