CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.5k downloads2y agoHugging Face02Emova-ollm /emova-sft-4m EMOVA-SFT-4M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.imageimage-to-text1M<n<10M6 likes3k downloads2y agoHugging Face03langswap /dialogs-ru-emotional-conversations Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational Russian speech, designed for dialog-oriented and emotional text-to-speech. Unlike existing Russian corpora — mostly single-speaker read speech or large but low-quality web-mined audio — Dialogs was recorded by professional theatre actors performing scripted dialogs face-to-face, capturing natural turn-taking, timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.audiotext-to-speechn<1K18 likes2.4k downloads2mo agoHugging Face04Emova-ollm /emova-sft-speech-231k EMOVA-SFT-Speech-231K 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-Speech-231K is a comprehensive dataset curated for omni-modal instruction tuning and emotional spoken dialogue. This dataset is created by converting existing text and visual instruction datasets via Text-to-Speech (TTS) tools. EMOVA-SFT-Speech-231K is part of EMOVA-Datasets collection and is used in… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-speech-231k.imageaudio-to-audio100K<n<1M3 likes313 downloads2y agoHugging Face05Emova-ollm /emova-asr-tts-eval EMOVA-ASR-TTS-Eval 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-ASR-TTS-Eval is a dataset designed for evaluating the ASR and TTS performance of Omni-modal LLMs. It is derived from the test-clean set of the LibriSpeech dataset. This dataset is part of the EMOVA-Datasets collection. We extract the speech units using the EMOVA Speech Tokenizer. Structure This… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-asr-tts-eval.textautomatic-speech-recognition1K<n<10K0 likes245 downloads2y agoHugging Face06datahiveai /arabic-multidialect-emotional-speech-demo DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request. Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.audioautomatic-speech-recognitionn<1K3 likes122 downloads5mo agoHugging Face07OmarAhmedSobhy /egyption-with-emotion-dataset Egption Text-Audio Dataset With Emotions and Diarization Creating datasets for TTS and ASR models with emotions and Diarization In case you want to focus only one speaker , you can fiter based on speaker_role Source Code if you want to collect more data from youtube, you can check this link 🙏 Acknowledgements This project makes use of the forced alignment model and Cohere ASR model provided by: MahmoudAshraf/mms-300m-1130-forced-aligner Cohere ASR Hubert… See the full description on the dataset page: https://huggingface.co/datasets/OmarAhmedSobhy/egyption-with-emotion-dataset.audioautomatic-speech-recognition1K<n<10K4 likes105 downloads5mo agoHugging Face08gaydmi /emo_ttsgated Tatar Dubbed Speech Sentence-level speech segments in Tatar, cut from Tatar-language dubs and aligned to their subtitles with CTC forced alignment. 7,257 sentences / 4:39:06 drawn from 11.85 h of source audio — dialogue is sparse in this material, so roughly 40% of the runtime is speech. Source source prefix Rows Duration Median similarity Берсерк, 25 episodes Берсерк - N серия 5,855 4:03:34 0.941 Мистер һәм миссис Смит Мистер һәм миссис Смит 1,402 0:35:32 0.909… See the full description on the dataset page: https://huggingface.co/datasets/gaydmi/emo_tts.audiotext-to-speech1K<n<10K0 likes40 downloads5d agoHugging Face09Emova-ollm /emova-sft-speech-eval EMOVA-SFT-Speech-Eval 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-Speech-Eval is an evaluation dataset curated for omni-modal instruction tuning and emotional spoken dialogue. This dataset is created by converting existing text and visual instruction datasets via Text-to-Speech (TTS) tools. EMOVA-SFT-Speech-Eval is part of EMOVA-Datasets collection, and the training… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-speech-eval.imageaudio-to-audio1K<n<10K1 likes38 downloads2y agoHugging Face10NathanRoll /eng-sports-radio-psst-iu-emotion-splits English Sports Radio Non-Neutral Emotion IU Splits Public non-neutral subset of NathanRoll/eng-sports-radio-psst-iu. Each row is one intonation unit with exactly three columns: audio: embedded 16 kHz mono audio for the IU text: a leading emotion special token followed by the Parakeet transcript accent: broadcast-location proxy accent label Neutral examples were removed. The remaining rows are split by emotion: joy: 247 rows, 0.287 audio hours surprise: 153 rows, 0.188 audio… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/eng-sports-radio-psst-iu-emotion-splits.audioautomatic-speech-recognitionn<1K0 likes33 downloads4mo agoHugging Face11sarthwa8 /indian-tts-emotion-60min indian-tts-emotion-60min A small, carefully curated text-to-speech dataset: ~68 minutes of clean, single-speaker-per-clip audio in Indian English (en-IN) and Hindi (hi-IN), sourced from YouTube, with accurate transcriptions and per-clip emotion/style tags. Built as a data-quality exercise: clips were filtered conservatively and a sample was verified by listening rather than shipped straight from an automated pipeline. Summary Language Clips Duration (min)… See the full description on the dataset page: https://huggingface.co/datasets/sarthwa8/indian-tts-emotion-60min.audiotext-to-speechn<1K1 likes20 downloads3mo agoHugging Face12somu9 /raw-emoceangated raw-emocean Large-scale English speech dataset for text-to-speech (TTS) model training. Designed for autoregressive TTS architectures (TADA, CSM, VALL-E style models). Dataset Summary Metric Value Parquet shards 7 Segment duration 3–8 seconds Sample rate 24,000 Hz (mono) ASR engine NVIDIA Parakeet TDT 0.6B v3 Format Parquet with embedded audio Dataset Schema Column Type Description audio Audio Waveform array + sampling… See the full description on the dataset page: https://huggingface.co/datasets/somu9/raw-emocean.audiotext-to-speech10K<n<100K1 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.