CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M271 likes15k downloads24d agoHugging Face02simon3000 /zenless-voice Zenless Voice Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero. Hugging Face 🤗 Zenless-Voice ModelScope Zenless-Voice Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index. Last update at 2026-09-17, game version 3.2.0 406720 wavs 78785 without speaker (19%) 123429 without transcription (30%) 83509 without inGameFilename (21%) Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.audioaudio-classification100K<n<1M4 likes6.3k downloads7d agoHugging Face03hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes5.4k downloads4y agoHugging Face04XRXRX /X-Voice-Dataset-Train X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling. Also the train set of X-Voice Model. Core Statistics Total Speech Duration: 420K hours 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.audiotext-to-speech10M<n<100M11 likes4.7k downloads5mo agoHugging Face05VoiceNet /emolia-thinking Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.audioaudio-classification100K<n<1M0 likes4.5k downloads2mo agoHugging Face06hezarai /common-voice-13-faThe Persian portion of the original CommonVoice 13 dataset at https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0 Load # Using HF Datasets from datasets import load_dataset dataset = load_dataset("hezarai/common-voice-13-fa", split="train") # Using Hezar from hezar.data import Dataset dataset = Dataset.load("hezarai/common-voice-13-fa", split="train") audioautomatic-speech-recognition10K<n<100K1 likes3k downloads2y agoHugging Face07zhifeixie /Voices-in-the-Wild-2M Voices in the Wild Project Page | Paper | GitHub Voices in the Wild (Voices-in-the-Wild-2M) is a large-scale automatic speech recognition (ASR) dataset designed for robustness training and evaluation under diverse, real-world acoustic conditions. It covers 7 classic acoustic phenomena (including noise, far-field speech, obstruction, echo/reverberation, recording artifacts, electronic distortion, and transmission dropout) and 54 physically plausible compound scenarios. The… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/Voices-in-the-Wild-2M.audioautomatic-speech-recognition50 likes2.9k downloads4mo agoHugging Face08kadirnar /voicehub-arena-seed-tts-eval VoiceHub Arena — full English Seed-TTS-Eval 35,904 synthesized WAV files: 33 model families × the same 1,088 target texts. The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB. All 198 shards and every WAV SHA256 were verified after backup. Interactive leaderboard and all audio samples · Source repository (access required). Contents audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.audiotext-to-speech10K<n<100K0 likes2.8k downloads8d agoHugging Face09simon3000 /starrail-voice StarRail Voice StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail. Hugging Face 🤗 StarRail-Voice ModelScope StarRail-Voice Last update at 2026-07-16, game version 4.4.0 403437 wavs 60164 without speaker (15%) 61375 without transcription (15%) 57869 without inGameFilename (14%) Dataset Details Dataset Description The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.audioaudio-classification100K<n<1M63 likes2.7k downloads2mo agoHugging Face10moaead /dialectal-arabic-voices Dialectal Arabic Voices An expanding collection of Arabic audio from YouTube, SoundCloud, and other sources. Currently labelled Palestinian Arabic (ps). 49,951 recordings · approximately 9,169.6 hours · 481.10 GB Column Description audio Original audio, embedded in the Parquet file transcript_text Empty for now; ASR transcripts will be added later language Dialect code: ps (Palestinian) source Original channel or account name Audio retains its original… See the full description on the dataset page: https://huggingface.co/datasets/moaead/dialectal-arabic-voices.audioautomatic-speech-recognition10K<n<100K0 likes1.7k downloads2h agoHugging Face11doduy1911 /voiceNam Dataset Tiếng Việt (Voice Nữ) Đây là bộ dữ liệu bao gồm file âm thanh và transcript tương ứng, được sử dụng cho việc train mô hình TTS (Text-to-Speech). Cấu trúc dữ liệu audio: File âm thanh (.wav) text: Nội dung văn bản tương ứng Cách sử dụng from datasets import load_dataset dataset = load_dataset("doduy1911/voiceNamNam", split="train") # Nghe thử mẫu đầu tiên print(dataset[0]["text"]) audiotext-to-speech1K<n<10K0 likes1.7k downloads8mo agoHugging Face12Muckylixx /mucky-voice-final Mucky Voice – Deutscher Einzelsprecher-Sprachdatensatz Ein kuratierter deutscher Sprachdatensatz mit Aufnahmen meiner eigenen Stimme. Der Datensatz enthält 11.817 geprüfte Audio-Text-Paare mit einer gesamten nutzbaren Sprachdauer von 43 Stunden, 43 Minuten und 37,1 Sekunden. Er eignet sich insbesondere für: deutsches Text-to-Speech-Training Sprecheranpassung und Voice Cloning automatische Spracherkennung Forschung zu spontaner Sprache Sprachbereinigung und Aussprachemodelle… See the full description on the dataset page: https://huggingface.co/datasets/Muckylixx/mucky-voice-final.audiotext-to-speech10K<n<100K0 likes1k downloads2mo agoHugging Face13hanamizuki-ai /genshin-voice-v3.5-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.5-mandarin.audiotext-to-speech10K<n<100K18 likes1k downloads3y agoHugging Face14dsfsi-anv /za-african-next-voicesgated Swivuriso: ZA-African Next Voices Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes. Dataset Paper: ArXiv - Work in Progress Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.audioautomatic-speech-recognition100K<n<1M16 likes1k downloads7mo agoHugging Face15VoiceArena /MonsoonASR-Open-ASR-leaderboard-en-IN Voice Arena Monsoon en-IN (public test) Part of the Open ASR Leaderboard, in the main board's default column set, so it contributes to the headline Average WER for every model listed. A conversational Indian English ASR test set that records who was speaking, not only what was said. Every clip carries twelve speaker attributes — gender, age, native district and state, education, occupation, income band, handset — so a difference between two systems can be traced to a group of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN.audioautomatic-speech-recognition1K<n<10K3 likes921 downloads26d agoHugging Face16Africanvoice /African_voices_naija 🇳🇬 WaZoBiaSpeech: 1,000+ Hour Nigerian Pidgin (pcm) Corpus Version: 30 Nov 2025 NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release. 🌍 Dataset Overview WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Nigerian Pidgin (pcm). This corpus is designed to accelerate the development of speech technology in African contexts, promoting… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_naija.audioautomatic-speech-recognition100K<n<1M0 likes857 downloads2d agoHugging Face17besimple-ai /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.audioautomatic-speech-recognitionn<1K14 likes729 downloads13d agoHugging Face18NandemoGHS /Japanese-Eroge-Voice-V2 Japanese-Eroge-Voice-V2 Description This is the successor to the Japanese-Eroge-Voice dataset. It consists of a significantly larger collection of audio-transcription pairs extracted from Japanese eroge (adult games). Note on Versioning: There is no overlap between this dataset (V2) and the previous version. All audio clips and transcriptions in V2 are distinct from those in the original version, providing entirely new data for research. This version (V2) expands the… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice-V2.audiotext-to-speech1M<n<10M53 likes677 downloads8mo agoHugging Face19JacobLinCool /common_voice_19_0_zh-TW Common Voice Corpus 19.0 Chinese (Taiwan) The test set is the same as the original test set, while validated_without_test includes all validated examples except those with sentence IDs that appear in the test set. validated_without_test has about 50,000 examples in total, equivalent to approximately 44 hours, and is intended for use as the training set. test has about 5,000 examples, which is approximately 5 hours. audioautomatic-speech-recognition10K<n<100K3 likes614 downloads2y agoHugging Face20starrydark /Kurisu_Voice Makise Kurisu Multilingual Voice Dataset 13,999 labelled clips (17.1423 hours) of Makise Kurisu, including the Amadeus Kurisu variant, in six languages, cut from the STEINS;GATE games, the anime, and three character songs. Every clip carries the transcript, a measured acoustic profile, an independent speaker-identity check, and — where it could be earned rather than guessed — an expressive tag. This is an unofficial, fan-made dataset with no affiliation to the STEINS;GATE rights… See the full description on the dataset page: https://huggingface.co/datasets/starrydark/Kurisu_Voice.audiotext-to-speech10K<n<100K0 likes609 downloads19d agoHugging Face21NandemoGHS /Japanese-Eroge-Voice Japanese-Eroge-Voice Description This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model. Preprocessing Steps The raw audio data has undergone the following preprocessing steps: Loudness Normalization: Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice.audiotext-to-speech100K<n<1M37 likes596 downloads1y agoHugging Face22language-and-voice-lab /samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.audioautomatic-speech-recognition10K<n<100K9 likes561 downloads3y agoHugging Face23unlimitedbytes /hailuo-ai-voices Hailuo AI Voices Dataset 🎤 A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks. 📊 Dataset Overview The dataset provides a comprehensive collection of voice samples with the following features: Feature Description Audio Files High-quality WAV format recordings Transcription Accurate transcriptions of each… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-voices.audiotext-to-speech10K<n<100K9 likes561 downloads2y agoHugging Face24srezas /farsi_voice_dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/srezas/farsi_voice_dataset.audioautomatic-speech-recognition100K<n<1M5 likes552 downloads2y agoHugging Face25MohamedRashad /common-voice-18-arabic Dataset Card for Common Voice 18 – Arabic Edition Dataset Summary This dataset is an unofficial Arabic-only extraction of Mozilla Common Voice Corpus 18.0, prepared for Automatic Speech Recognition (ASR) research and development. It is derived from the original Common Voice 18 release and filtered to include Arabic (ar) speech data only, while preserving the original dataset structure, splits, and metadata fields. The dataset consists of validated, unvalidated, and… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/common-voice-18-arabic.audioautomatic-speech-recognition100K<n<1M5 likes483 downloads9mo agoHugging Face26ghananlpcommunity /ghana-one-voice Ghana One Voice Speech in 43 Ghanaian and West African languages, all converted into a single consistent voice. Every clip keeps its original transcript, so the dataset pairs one speaker's voice with the phonetic range of dozens of languages. Roughly 5 hours per language, about 215 hours in total. What this is Source audio comes from many different speakers, recording conditions and microphones. Each clip has been passed through ghana-vc, a voice-conversion model… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-one-voice.audiotext-to-speech100K<n<1M0 likes437 downloads7d agoHugging Face27deepghs /arknights_voices_zh ZH Voice-Text Dataset for Arknights Waifus This is the ZH voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models. Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset. 12431 records, 25.9 hours in total. Average duration is approximately 7.49s. id char_id voice_actor_name voice_title voice_text time sample_rate file_size filename mimetype file_url char_106_franka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_zh.tabularautomatic-speech-recognition10K<n<100K6 likes431 downloads2y agoHugging Face28ultemica /genshin-impact-voices Genshin Impact — Voice Lines (Multi-Language) An archive of character voice data extracted from Genshin Impact (原神), repackaged as Parquet shards per audio language. Dataset Summary Field Value Game Genshin Impact (原神) Publisher HoYoverse / miHoYo Co., Ltd. Languages 中文 (zh), 日本語 (ja), English (en), 한국어 (ko) Game version 6.3 Source format WAV + sidecar transcripts (.lab / .txt) Distribution format Apache Parquet (zstd), ~500 MiB audio per shard… See the full description on the dataset page: https://huggingface.co/datasets/ultemica/genshin-impact-voices.audioautomatic-speech-recognition100K<n<1M0 likes399 downloads5mo agoHugging Face29VoiceNet /emolia emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired) This is the emolia-balanced-5M-subset corpus repackaged for high-quality audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz (PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json samples. The JSON sidecar carries the full annotation stack: Original metadata (id, text, duration, speaker, language, dnsmos). A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.audioaudio-classification1M<n<10M1 likes385 downloads5mo agoHugging Face30xmodar /commonvoice-12.0-arabic-voice-converted Dataset Card for Voice Converted Arabic Common Voice 12.0 This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.audioautomatic-speech-recognition100K<n<1M8 likes358 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.