CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fixie-ai /common_voice_17_0audio10M<n<100M18 likes202k downloads2y agoHugging Face02VoiceHub /voicehub-arena-seed-tts-eval VoiceHub Arena — native TTS evaluations Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA speaker SIM and UTMOS22 measurements. The full campaign is still running. Each generation method is evaluated separately using its publisher's native API. Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target diagnostic pilots are stored separately and must not be treated as full scores. Interactive demo · Source code Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.audiotext-to-speech0 likes25k downloads5d agoHugging Face03simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M271 likes15k downloads24d agoHugging Face04hlt-lab /voicebench License The dataset is available under the Apache 2.0 license. Citation If you use the VoiceBench dataset in your research, please cite the following paper: @article{chen2024voicebench, title={VoiceBench: Benchmarking LLM-Based Voice Assistants}, author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou}, journal={arXiv preprint arXiv:2410.17196}, year={2024} } audio10K<n<100K16 likes7.6k downloads1y agoHugging Face05simon3000 /zenless-voice Zenless Voice Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero. Hugging Face 🤗 Zenless-Voice ModelScope Zenless-Voice Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index. Last update at 2026-09-17, game version 3.2.0 406720 wavs 78785 without speaker (19%) 123429 without transcription (30%) 83509 without inGameFilename (21%) Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.audioaudio-classification100K<n<1M4 likes6.3k downloads7d agoHugging Face06AquaV /fallout-4-voicesaudio10K<n<100K3 likes5.5k downloads2y agoHugging Face07hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes5.4k downloads4y agoHugging Face08pedrohavay /portuguese-male-voice-A-datasetaudio4 likes5.3k downloads3y agoHugging Face09krutrim-ai-labs /VoiceAgentBench VoiceAgentBench This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978). VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.audio1K<n<10K9 likes4.9k downloads7mo agoHugging Face10ayf3 /numberblocks-one-voice-datasetaudio1K<n<10K7 likes4.8k downloads5d agoHugging Face11XRXRX /X-Voice-Dataset-Train X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling. Also the train set of X-Voice Model. Core Statistics Total Speech Duration: 420K hours 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.audiotext-to-speech10M<n<100M11 likes4.7k downloads5mo agoHugging Face12VoiceNet /emolia-thinking Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.audioaudio-classification100K<n<1M0 likes4.5k downloads2mo agoHugging Face13gpt-omni /VoiceAssistant-400Kaudio100K<n<1M100 likes4.3k downloads2y agoHugging Face14UncovAI /Real_Voiceaudio100K<n<1M1 likes4.2k downloads1y agoHugging Face15anke01 /uyghur-common-voice-tts Uyghur Common Voice TTS Dataset A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice. Dataset Summary Property Value Language Uyghur (ug) Total Samples 43,054 Train Samples 40,901 Validation Samples 2,153 Audio Format WAV Source Mozilla Common Voice License CC0-1.0 Dataset Structure / ├── train.jsonl # Training data (40,901 samples) ├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.audiotext-to-speech10K<n<100K0 likes3.9k downloads7mo agoHugging Face16SpeechTest /common_voice_16_0audio100K<n<1M0 likes3.5k downloads8mo agoHugging Face17VoiceOfML /Japanese-Materials 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/Japanese-Materials/discussions 提出。 此仓库存储日共资料:https://huggingface.co/datasets/VoiceOfML/Japanese-Materials/tree/main 。 请使用:https://voiceofml-search.hf.space/Japanese-Materials 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/Japanese-Materials )。 可使用:https://voiceofml-search.hf.space/Japanese-Materials?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/Japanese-Materials?wide=1 )。… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/Japanese-Materials.audion<1K0 likes3k downloads11d agoHugging Face18hezarai /common-voice-13-faThe Persian portion of the original CommonVoice 13 dataset at https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0 Load # Using HF Datasets from datasets import load_dataset dataset = load_dataset("hezarai/common-voice-13-fa", split="train") # Using Hezar from hezar.data import Dataset dataset = Dataset.load("hezarai/common-voice-13-fa", split="train") audioautomatic-speech-recognition10K<n<100K1 likes3k downloads2y agoHugging Face19zhifeixie /Voices-in-the-Wild-2M Voices in the Wild Project Page | Paper | GitHub Voices in the Wild (Voices-in-the-Wild-2M) is a large-scale automatic speech recognition (ASR) dataset designed for robustness training and evaluation under diverse, real-world acoustic conditions. It covers 7 classic acoustic phenomena (including noise, far-field speech, obstruction, echo/reverberation, recording artifacts, electronic distortion, and transmission dropout) and 54 physically plausible compound scenarios. The… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/Voices-in-the-Wild-2M.audioautomatic-speech-recognition50 likes2.9k downloads4mo agoHugging Face20kadirnar /voicehub-arena-seed-tts-eval VoiceHub Arena — full English Seed-TTS-Eval 35,904 synthesized WAV files: 33 model families × the same 1,088 target texts. The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB. All 198 shards and every WAV SHA256 were verified after backup. Interactive leaderboard and all audio samples · Source repository (access required). Contents audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.audiotext-to-speech10K<n<100K0 likes2.8k downloads8d agoHugging Face21simon3000 /starrail-voice StarRail Voice StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail. Hugging Face 🤗 StarRail-Voice ModelScope StarRail-Voice Last update at 2026-07-16, game version 4.4.0 403437 wavs 60164 without speaker (15%) 61375 without transcription (15%) 57869 without inGameFilename (14%) Dataset Details Dataset Description The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.audioaudio-classification100K<n<1M63 likes2.7k downloads2mo agoHugging Face22vinhn8n /voicedaolyaudio1K<n<10K0 likes2.6k downloads20d agoHugging Face23laion /laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants. Updated Composition Voices and Languages English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.audio1M<n<10M6 likes2.6k downloads1y agoHugging Face24mteb /common_voice_21_0_miniaudio10K<n<100K0 likes2.1k downloads9mo agoHugging Face25JacobLinCool /VoiceBank-DEMAND-16kaudioaudio-to-audio10K<n<100K17 likes2.1k downloads2y agoHugging Face26kanuli1983 /japanese-listening-voicevox-backupaudio0 likes2k downloads18d agoHugging Face27goea /arc-voicesamples-generatedaudio100K<n<1M1 likes2k downloads3h agoHugging Face28VoiceOfML /Omnibus 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/Omnibus/discussions 提出。 此仓库存储书纳百川专题:https://huggingface.co/datasets/VoiceOfML/Omnibus/tree/main 。 请使用:https://voiceofml-search.hf.space/Omnibus 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/Omnibus )。 可使用:https://voiceofml-search.hf.space/Omnibus?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/Omnibus?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone without large files - just their… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/Omnibus.audion<1K0 likes1.9k downloads21d agoHugging Face29BertilBraun /voice-light-synthetic-audio Voice-Light Synthetic Audio English-only synthetic conversational speech for training and evaluating streaming turn-taking models. The corpus focuses on end-of-turn prediction, continuation holds, short backchannels, interruptions, and response timing. The dataset contains user-side FLAC speech units plus typed conversation plans, rendering provenance, quality ledgers, and deterministic reconstruction metadata. Assistant speech is represented as a time-varying… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-synthetic-audio.audio10K<n<100K0 likes1.9k downloads27d agoHugging Face30lmms-lab-audio /voicebench License The dataset is available under the Apache 2.0 license. Citation If you use the VoiceBench dataset in your research, please cite the following paper: @article{chen2024voicebench, title={VoiceBench: Benchmarking LLM-Based Voice Assistants}, author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou}, journal={arXiv preprint arXiv:2410.17196}, year={2024} } audio10K<n<100K1 likes1.9k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.