CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AquaV /genshin-voices-separated21 likes233k downloads2y agoHugging Face02fixie-ai /common_voice_17_0audio10M<n<100M18 likes194k downloads2y agoHugging Face03VoiceOfML /VOMEBOOK 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/VOMEBOOK/discussions 提出。 此仓库存储马列之声电子书:https://huggingface.co/datasets/VoiceOfML/VOMEBOOK/tree/main 。 请使用:https://voiceofml-search.hf.space/VOMEBOOK 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/VOMEBOOK )。 可使用:https://voiceofml-search.hf.space/VOMEBOOK?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/VOMEBOOK?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone without large files - just their… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/VOMEBOOK.document10K<n<100K8 likes63k downloads8d agoHugging Face04fsicoli /common_voice_15_0 Dataset Card for Common Voice Corpus 15.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 15. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_15_0.automatic-speech-recognition100B<n<1T6 likes44k downloads3y agoHugging Face05fsicoli /common_voice_22_0 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_22_0.automatic-speech-recognition100B<n<1T20 likes26k downloads1y agoHugging Face06VoiceHub /voicehub-arena-seed-tts-eval VoiceHub Arena — native TTS evaluations Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA speaker SIM and UTMOS22 measurements. The full campaign is still running. Each generation method is evaluated separately using its publisher's native API. Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target diagnostic pilots are stored separately and must not be treated as full scores. Interactive demo · Source code Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.audiotext-to-speech0 likes25k downloads2d agoHugging Face07legacy-datasets /common_voiceCommon Voice is Mozilla's initiative to help teach machines how real people speak. The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.automatic-speech-recognition100K<n<1M148 likes15k downloads2y agoHugging Face08fsicoli /common_voice_17_0 Dataset Card for Common Voice Corpus 17.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_17_0.automatic-speech-recognition100B<n<1T19 likes14k downloads2y agoHugging Face09simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M271 likes14k downloads21d agoHugging Face10fsicoli /common_voice_16_0 Dataset Card for Common Voice Corpus 16.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 16. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_16_0.automatic-speech-recognition100B<n<1T4 likes13k downloads3y agoHugging Face11VoiceOfML /Teachers 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/Teachers/discussions 提出。 此仓库存储导师著作:https://huggingface.co/datasets/VoiceOfML/Teachers/tree/main 。 请使用:https://voiceofml-search.hf.space/Teachers 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/Teachers )。 可使用:https://voiceofml-search.hf.space/Teachers?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/Teachers?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone without large files - just their… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/Teachers.document1K<n<10K1 likes12k downloads8d agoHugging Face12ebook2audiobook /E2A-Voices0 likes10k downloads10d agoHugging Face13hlt-lab /voicebench License The dataset is available under the Apache 2.0 license. Citation If you use the VoiceBench dataset in your research, please cite the following paper: @article{chen2024voicebench, title={VoiceBench: Benchmarking LLM-Based Voice Assistants}, author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou}, journal={arXiv preprint arXiv:2410.17196}, year={2024} } audio10K<n<100K16 likes7.6k downloads1y agoHugging Face14VoiceOfML /MLMRL-Library 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/MLMRL-Library/discussions 提出。 此仓库存储重要书库备份:https://huggingface.co/datasets/VoiceOfML/MLMRL-Library/tree/main 。 请使用:https://voiceofml-search.hf.space/MLMRL-Library 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/MLMRL-Library )。 可使用:https://voiceofml-search.hf.space/MLMRL-Library?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/MLMRL-Library?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/MLMRL-Library.1 likes6.7k downloads8d agoHugging Face15mteb /common_voice_21_00 likes6.1k downloads1y agoHugging Face16anke01 /uyghur-common-voice-tts Uyghur Common Voice TTS Dataset A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice. Dataset Summary Property Value Language Uyghur (ug) Total Samples 43,054 Train Samples 40,901 Validation Samples 2,153 Audio Format WAV Source Mozilla Common Voice License CC0-1.0 Dataset Structure / ├── train.jsonl # Training data (40,901 samples) ├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.audiotext-to-speech10K<n<100K0 likes6k downloads6mo agoHugging Face17VoiceOfML /SovMaterials 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/SovMaterials/discussions 提出。 此仓库存储苏联资料:https://huggingface.co/datasets/VoiceOfML/SovMaterials/tree/main 。 请使用:https://voiceofml-search.hf.space/SovMaterials 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/SovMaterials )。 可使用:https://voiceofml-search.hf.space/SovMaterials?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/SovMaterials?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone without… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/SovMaterials.document1K<n<10K5 likes5.9k downloads8d agoHugging Face18VoiceOfML /MLMRL-Hub 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/MLMRL-Hub/discussions 提出。 此仓库存储马列毛主义与革命左翼仓储中心和图书馆的未重复资料(仅经一次md5检测,压缩包内容未去重。):https://huggingface.co/datasets/VoiceOfML/MLMRL-Hub/tree/main 。 请使用https://voiceofml-search.hf.space/MLMRL-Hub 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/MLMRL-Hub )。 可使用:https://voiceofml-search.hf.space/MLMRL-Hub?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/MLMRL-Hub?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/MLMRL-Hub.0 likes5.8k downloads8d agoHugging Face19pedrohavay /portuguese-male-voice-A-datasetaudio4 likes5.3k downloads3y agoHugging Face20hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes5.2k downloads4y agoHugging Face21krutrim-ai-labs /VoiceAgentBench VoiceAgentBench This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978). VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.audio1K<n<10K9 likes5.1k downloads7mo agoHugging Face22XRXRX /X-Voice-Dataset-Train X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling. Also the train set of X-Voice Model. Core Statistics Total Speech Duration: 420K hours 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.audiotext-to-speech10M<n<100M11 likes4.8k downloads5mo agoHugging Face23ayf3 /numberblocks-one-voice-datasetaudio1K<n<10K7 likes4.8k downloads2d agoHugging Face24UncovAI /Real_Voiceaudio100K<n<1M1 likes4.3k downloads1y agoHugging Face25gpt-omni /VoiceAssistant-400Kaudio100K<n<1M100 likes4.2k downloads2y agoHugging Face26VoiceNet /emolia-thinking Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.audioaudio-classification100K<n<1M0 likes4.2k downloads2mo agoHugging Face27AquaV /fallout-4-voicesaudio10K<n<100K3 likes4.1k downloads2y agoHugging Face28Nihilux /common_voice_speak_text0 likes4.1k downloads16d agoHugging Face29VoiceOfML /GPCREducation 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/GPCREducation/discussions 提出。 此仓库存储文化大革命教材:https://huggingface.co/datasets/VoiceOfML/GPCREducation/tree/main 。 请使用:https://voiceofml-search.hf.space/GPCREducation 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/GPCREducation )。 可使用:https://voiceofml-search.hf.space/GPCREducation?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/GPCREducation?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/GPCREducation.document1K<n<10K1 likes3.8k downloads3mo agoHugging Face30SpeechTest /common_voice_16_0audio100K<n<1M0 likes3.6k downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.