CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M271 likes15k downloads26d agoHugging Face02CoRal-project /coral-v3gated CoRal: Danish Conversational and Read-aloud Dataset Version 3.0 Dataset Overview CoRal is a comprehensive Automatic Speech Recognition (ASR) dataset designed to capture the diversity of the Danish language across various dialects, accents, genders, and age groups. The primary goal of the CoRal dataset is to provide a robust resource for training and evaluating ASR models that can understand and transcribe spoken Danish in all its variations. Key Features… See the full description on the dataset page: https://huggingface.co/datasets/CoRal-project/coral-v3.audioautomatic-speech-recognition100K<n<1M6 likes9.9k downloads7mo agoHugging Face03amithm3 /shrutilipiaudioautomatic-speech-recognition1M<n<10M6 likes7k downloads2y agoHugging Face04simon3000 /zenless-voice Zenless Voice Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero. Hugging Face 🤗 Zenless-Voice ModelScope Zenless-Voice Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index. Last update at 2026-09-17, game version 3.2.0 406720 wavs 78785 without speaker (19%) 123429 without transcription (30%) 83509 without inGameFilename (21%) Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.audioaudio-classification100K<n<1M5 likes6.4k downloads9d agoHugging Face05hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes5.5k downloads4y agoHugging Face06meldynamics /liepa-3 LIEPA-3 — Lithuanian Speech Corpus Didysis lietuvių kalbos garsynas (LIEPA-3) Dataset Summary LIEPA-3 is a large, open corpus of Lithuanian speech (~10,000 hours, ~7.5 million audio files) built for automatic speech recognition (ASR), text-to-speech (TTS) and linguistic research. It spans read, spontaneous, phonetically-annotated and dialectal speech recorded under a wide range of conditions (studio, dictaphone, radio, TV, telephone, audiobooks). Official… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-3.audioautomatic-speech-recognition1M<n<10M4 likes3.3k downloads3mo agoHugging Face07simon3000 /starrail-voice StarRail Voice StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail. Hugging Face 🤗 StarRail-Voice ModelScope StarRail-Voice Last update at 2026-07-16, game version 4.4.0 403437 wavs 60164 without speaker (15%) 61375 without transcription (15%) 57869 without inGameFilename (14%) Dataset Details Dataset Description The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.audioaudio-classification100K<n<1M63 likes2.7k downloads2mo agoHugging Face08Shirali /ISSAI_KSC_335RS_v_1_1 Dataset Card for "ISSAI_KSC_335RS_v_1_1" Kazakh Speech Corpus (KSC) Identifier: SLR102 Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours) Category: Speech License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN] About this resource: A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.audioautomatic-speech-recognition100K<n<1M3 likes2.5k downloads4y agoHugging Face09espnet /yodas3 YODAS v3 Paper YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 and v2 versions of YODAS, to guarantee that there are no overlaps in the data. For… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas3.audio-to-audio2 likes2.4k downloads1d agoHugging Face10Haitam03 /warsh-segments-v3 Haitam03/warsh-v3 Warsh (Rewayat Warsh A'n Nafi') Quran recitation, segmented at waqf with obadx/recitation-segmenter-v2. Built with warsh-data. Layout path what data/<reciter>/<surah>.parquet one file per source recording, audio embedded as 16 kHz mono FLAC raw/<reciter>/<surah>.mp3 the source recording it came from segment_params.json the settings this corpus was produced with One parquet per source recording, named after it, so re-running a… See the full description on the dataset page: https://huggingface.co/datasets/Haitam03/warsh-segments-v3.audioautomatic-speech-recognition100K<n<1M0 likes2k downloads29d agoHugging Face11MushanW /GLOBE_V3 Important notice Differences between V3 version and two previous versions (V1|V2): This version is built base on Common Voice 21.0 English Subset. This version only includes utterance that are an exact match with the transcription from Whisper V3 LARGE (CER == 0). This version includes the original Common Voice metadata (age, gender, accent, and ID). All audio files in this version are at 24kHz sampling rate. All audio files in this version are unenhanced. (We’d greatly… See the full description on the dataset page: https://huggingface.co/datasets/MushanW/GLOBE_V3.audiotext-to-audio100K<n<1M2 likes1.8k downloads1y agoHugging Face12mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3 Pseudolabel Malaysian Youtube videos using Whisper Large V3 Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper Each audio is 30 seconds. Each audio saved in 16k sample rate. audioautomatic-speech-recognition3 likes1.2k downloads3y agoHugging Face13oddadmix /lahgtna-v3-small Lahgtna — Dialect-Balanced Arabic ASR (v3 small) A dialect-balanced multi-dialect Arabic speech-recognition corpus: 54,600 clips / 267.3 hours across 13 Arabic dialects, 16 kHz mono. Each dialect is evenly represented — 4,000 train + 200 test clips per dialect — so models and evaluations aren't skewed toward high-resource dialects (e.g. Egyptian/Gulf). Used to train the oddadmix v2 dialectal-ASR model family. Duration by dialect Dialect Train (h) Train clips… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/lahgtna-v3-small.audioautomatic-speech-recognition10K<n<100K7 likes1.2k downloads2mo agoHugging Face14hanamizuki-ai /genshin-voice-v3.5-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.5-mandarin.audiotext-to-speech10K<n<100K18 likes1k downloads3y agoHugging Face15TheNHz /ellipsis-lrs3-rawgated LRS3-TED — verified mirror A mirror of the LRS3-TED dataset (Lip Reading Sentences 3), preserved because the official distribution has been discontinued. This repository adds no new data: it is a re-hosted copy with a full verification report against the official file list, so you know exactly what is and is not here. Attribution LRS3-TED was created by Triantafyllos Afouras, Joon Son Chung and Andrew Zisserman (Visual Geometry Group, University of Oxford): T.… See the full description on the dataset page: https://huggingface.co/datasets/TheNHz/ellipsis-lrs3-raw.textautomatic-speech-recognition1K<n<10K8 likes920 downloads2mo agoHugging Face16kennethli319 /seamless-interaction-jefferson-annotations Seamless Interaction Jefferson-Style Annotations An automatic, turn-oriented annotation layer for the Meta Seamless Interaction Dataset. It compares the dataset's traditional transcript with an ASR-derived Jefferson-style condition and supplies speech-act, communicative-purpose, interactional-signal, alignment, and quality fields. This is a derived noncommercial research dataset. It does not redistribute the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.tabularautomatic-speech-recognition100K<n<1M0 likes881 downloads2mo agoHugging Face17suleiman2003 /afri-temp-data3 AfricanVoices Hausa -- Train Split from datasets import load_dataset ds = load_dataset("suleiman2003/afri-temp-data3", split="train") print(ds[0]) audioautomatic-speech-recognition1K<n<10K0 likes878 downloads2mo agoHugging Face18i4tech /liepa3 LIEPA-3 Lithuanian Speech Corpus This repository repackages the original LIEPA-3 release into Hugging Face Parquet shards with embedded FLAC audio bytes. The original transcriptions are kept as released: normalized lowercase Lithuanian text without punctuation, digits, capitalization, or other symbols. Recommended use: read: cleanest subset and the default starting point for TTS or ASR. spon: spontaneous/broadcast/media speech; useful for ASR, not a clean TTS default. dial:… See the full description on the dataset page: https://huggingface.co/datasets/i4tech/liepa3.audioautomatic-speech-recognition1M<n<10M0 likes833 downloads3mo agoHugging Face19laion /voice-acting-edge-top3 Edge-case reward-Top-3 — voice annotations 1,494,503 synthetic expressive-speech utterances (the reward-Top-3 selection of the edge-case corpus), annotated with: a CrisperWhisper-format transcript with corrected vocal bursts — each surviving burst carries a class name from laion/vocal-burst-detector-v2 at its original timestamp; raw voice scores at two granularities — whole utterance and per sentence — from laion/Empathic-Insight-Voice-Plus (all 40 emotions), the 57 VoiceNet… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-edge-top3.automatic-speech-recognition1M<n<10M1 likes819 downloads14d agoHugging Face20notmax123 /ivirits-audio-v2-30s ivrit.ai audio-v2 — 2–30 s segments ivrit-ai/audio-v2 (>20k hours of Hebrew audio) cut into 2–30 second speech segments with machine transcripts, ready for ASR fine-tuning. How it was built VAD — Silero VAD (ONNX) over each episode decoded to 16 kHz mono. Speech regions longer than 30 s are split at the quietest sufficiently-long pause inside the window, so cuts land in silence rather than mid-word. Regions shorter than 2 s are dropped. Transcription —… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/ivirits-audio-v2-30s.audioautomatic-speech-recognition1M<n<10M0 likes775 downloads2mo agoHugging Face21Reza2kn /nasle-mana-clean-chunked-30s Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk. Configuration Rows Columns Meaning labeled (train/) 9,886 audio, label Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.audioautomatic-speech-recognition1K<n<10K0 likes697 downloads27d agoHugging Face22Reza2kn /nasle-mana-clean-chunked-30s-avasanj Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here. Splits Split Rows Audio Columns labeled 4,981 41.41 hours audio, label to_transcribe 11,127 92.72 hours audio The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.audioautomatic-speech-recognition10K<n<100K1 likes676 downloads19d agoHugging Face23NathanRoll /global-news-radio-30s Global News Radio Dataset Multilingual news radio recordings from 51 languages across 42 countries. Recordings 51 Total audio 1500 min (25.0 h) Format MP3 16kHz mono 64kbps Parquet shards 11 Languages 51 Countries 42 Size 687 MB Languages Amharic, Arabic, Bashkir, Basque, Belarusian, Bengali, Brazilian Portuguese,Portugues Do Brasil,Português Brasil, Catalan, Croatian, Czech, Danish, Dutch, English, Estonian, Faroese, Finnish, Flemish… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-30s.audioautomatic-speech-recognitionn<1K0 likes637 downloads6mo agoHugging Face24WhissleAI /Meta_STT_ZH_AIShell3 Meta Speech Recognition Mandarin Dataset (AISHELL3) This dataset contains both metadata and audio files for Mandarin speech recognition samples from the AISHELL3 corpus. Dataset Statistics Splits and Sample Counts train: 60098 samples valid: 3163 samples test: 24772 samples Example Samples train { "audio_filepath": "/external4/datasets/Mandarin/AISHELL3/wavs_train/SSB00430356.wav", "text": "她以 ENTITY_PRODUCT 滴鸡精 END 调养身体。 AGE_14_25… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_ZH_AIShell3.audioautomatic-speech-recognition10K<n<100K0 likes443 downloads1y agoHugging Face25laion /voice-acting-reinterpretations-top3 Voice-acting reinterpretations — reward-ranked top 3 of 64 ~20,000 acting prompts, each re-performed 64 times by laion/moss-tts-local-transformer-4.55b-voice-acting-v2, with the three highest-reward takes published here — plus the original clip each one reinterprets. Roughly 80,000 audio files (~20k originals + ~60k takes). The companion release …-raw64 keeps all 64 candidates per group, so the full reward distribution — not just its upper tail — stays available. ⚠️… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-reinterpretations-top3.text-to-speech10K<n<100K0 likes427 downloads14d agoHugging Face26projecte-aina /parlament_parla_v3 Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.audioautomatic-speech-recognition100K<n<1M1 likes418 downloads2y agoHugging Face27novelwolde36 /WaxalNLP Waxal Datasets The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus. Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/novelwolde36/WaxalNLP.audioautomatic-speech-recognition1M<n<10M1 likes413 downloads7mo agoHugging Face28ai-music4you3 /enhanced-audiosnippets-long-2-8M Enhanced Audiosnippets Long 2.8M Enhanced version of mitermix/audiosnippets_long_2_8M with speech enhancement, emotion annotations, speaker embeddings, and comprehensive metadata analysis. Dataset Summary Metric Value Total samples 2,633,037 Total audio hours 4,932 h Duration range 3.0s - 1124.3s Mean duration 6.7s Audio format WAV, 48kHz mono Tar files 1,410 Processing Pipeline Each audio sample was processed through: Speech… See the full description on the dataset page: https://huggingface.co/datasets/ai-music4you3/enhanced-audiosnippets-long-2-8M.tabularaudio-classification1M<n<10M1 likes404 downloads6mo agoHugging Face29KeisukeMiyamoto /nhk-archive-audio-30sgated NHK Archives Audio 30s This is a Japanese speech corpus derived from NHK Archives Audio. Audio from public NHK Archives records was segmented into clips of up to 30 seconds using voice activity detection. The dataset contains 344,722 accepted clips, totaling 1,661.01 hours. Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text contains LLM-assisted corrections based on the transcript and available source title and description. This… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio-30s.audioautomatic-speech-recognition100K<n<1M0 likes371 downloads18h agoHugging Face30hanamizuki-ai /genshin-voice-v3.4-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.4-mandarin.audiotext-to-speech10K<n<100K11 likes344 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.