CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M271 likes15k downloads26d agoHugging Face02CoRal-project /coral-v3gated CoRal: Danish Conversational and Read-aloud Dataset Version 3.0 Dataset Overview CoRal is a comprehensive Automatic Speech Recognition (ASR) dataset designed to capture the diversity of the Danish language across various dialects, accents, genders, and age groups. The primary goal of the CoRal dataset is to provide a robust resource for training and evaluating ASR models that can understand and transcribe spoken Danish in all its variations. Key Features… See the full description on the dataset page: https://huggingface.co/datasets/CoRal-project/coral-v3.audioautomatic-speech-recognition100K<n<1M6 likes9.9k downloads7mo agoHugging Face03amithm3 /shrutilipiaudioautomatic-speech-recognition1M<n<10M6 likes7k downloads2y agoHugging Face04simon3000 /zenless-voice Zenless Voice Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero. Hugging Face 🤗 Zenless-Voice ModelScope Zenless-Voice Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index. Last update at 2026-09-17, game version 3.2.0 406720 wavs 78785 without speaker (19%) 123429 without transcription (30%) 83509 without inGameFilename (21%) Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.audioaudio-classification100K<n<1M5 likes6.4k downloads9d agoHugging Face05hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes5.5k downloads4y agoHugging Face06meldynamics /liepa-3 LIEPA-3 — Lithuanian Speech Corpus Didysis lietuvių kalbos garsynas (LIEPA-3) Dataset Summary LIEPA-3 is a large, open corpus of Lithuanian speech (~10,000 hours, ~7.5 million audio files) built for automatic speech recognition (ASR), text-to-speech (TTS) and linguistic research. It spans read, spontaneous, phonetically-annotated and dialectal speech recorded under a wide range of conditions (studio, dictaphone, radio, TV, telephone, audiobooks). Official… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-3.audioautomatic-speech-recognition1M<n<10M4 likes3.3k downloads3mo agoHugging Face07simon3000 /starrail-voice StarRail Voice StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail. Hugging Face 🤗 StarRail-Voice ModelScope StarRail-Voice Last update at 2026-07-16, game version 4.4.0 403437 wavs 60164 without speaker (15%) 61375 without transcription (15%) 57869 without inGameFilename (14%) Dataset Details Dataset Description The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.audioaudio-classification100K<n<1M63 likes2.7k downloads2mo agoHugging Face08Shirali /ISSAI_KSC_335RS_v_1_1 Dataset Card for "ISSAI_KSC_335RS_v_1_1" Kazakh Speech Corpus (KSC) Identifier: SLR102 Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours) Category: Speech License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN] About this resource: A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.audioautomatic-speech-recognition100K<n<1M3 likes2.5k downloads4y agoHugging Face09Haitam03 /warsh-segments-v3 Haitam03/warsh-v3 Warsh (Rewayat Warsh A'n Nafi') Quran recitation, segmented at waqf with obadx/recitation-segmenter-v2. Built with warsh-data. Layout path what data/<reciter>/<surah>.parquet one file per source recording, audio embedded as 16 kHz mono FLAC raw/<reciter>/<surah>.mp3 the source recording it came from segment_params.json the settings this corpus was produced with One parquet per source recording, named after it, so re-running a… See the full description on the dataset page: https://huggingface.co/datasets/Haitam03/warsh-segments-v3.audioautomatic-speech-recognition100K<n<1M0 likes2k downloads29d agoHugging Face10MushanW /GLOBE_V3 Important notice Differences between V3 version and two previous versions (V1|V2): This version is built base on Common Voice 21.0 English Subset. This version only includes utterance that are an exact match with the transcription from Whisper V3 LARGE (CER == 0). This version includes the original Common Voice metadata (age, gender, accent, and ID). All audio files in this version are at 24kHz sampling rate. All audio files in this version are unenhanced. (We’d greatly… See the full description on the dataset page: https://huggingface.co/datasets/MushanW/GLOBE_V3.audiotext-to-audio100K<n<1M2 likes1.8k downloads1y agoHugging Face11mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3 Pseudolabel Malaysian Youtube videos using Whisper Large V3 Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper Each audio is 30 seconds. Each audio saved in 16k sample rate. audioautomatic-speech-recognition3 likes1.2k downloads3y agoHugging Face12oddadmix /lahgtna-v3-small Lahgtna — Dialect-Balanced Arabic ASR (v3 small) A dialect-balanced multi-dialect Arabic speech-recognition corpus: 54,600 clips / 267.3 hours across 13 Arabic dialects, 16 kHz mono. Each dialect is evenly represented — 4,000 train + 200 test clips per dialect — so models and evaluations aren't skewed toward high-resource dialects (e.g. Egyptian/Gulf). Used to train the oddadmix v2 dialectal-ASR model family. Duration by dialect Dialect Train (h) Train clips… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/lahgtna-v3-small.audioautomatic-speech-recognition10K<n<100K7 likes1.2k downloads2mo agoHugging Face13hanamizuki-ai /genshin-voice-v3.5-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.5-mandarin.audiotext-to-speech10K<n<100K18 likes1k downloads3y agoHugging Face14kennethli319 /seamless-interaction-jefferson-annotations Seamless Interaction Jefferson-Style Annotations An automatic, turn-oriented annotation layer for the Meta Seamless Interaction Dataset. It compares the dataset's traditional transcript with an ASR-derived Jefferson-style condition and supplies speech-act, communicative-purpose, interactional-signal, alignment, and quality fields. This is a derived noncommercial research dataset. It does not redistribute the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.tabularautomatic-speech-recognition100K<n<1M0 likes881 downloads2mo agoHugging Face15suleiman2003 /afri-temp-data3 AfricanVoices Hausa -- Train Split from datasets import load_dataset ds = load_dataset("suleiman2003/afri-temp-data3", split="train") print(ds[0]) audioautomatic-speech-recognition1K<n<10K0 likes878 downloads2mo agoHugging Face16i4tech /liepa3 LIEPA-3 Lithuanian Speech Corpus This repository repackages the original LIEPA-3 release into Hugging Face Parquet shards with embedded FLAC audio bytes. The original transcriptions are kept as released: normalized lowercase Lithuanian text without punctuation, digits, capitalization, or other symbols. Recommended use: read: cleanest subset and the default starting point for TTS or ASR. spon: spontaneous/broadcast/media speech; useful for ASR, not a clean TTS default. dial:… See the full description on the dataset page: https://huggingface.co/datasets/i4tech/liepa3.audioautomatic-speech-recognition1M<n<10M0 likes833 downloads3mo agoHugging Face17notmax123 /ivirits-audio-v2-30s ivrit.ai audio-v2 — 2–30 s segments ivrit-ai/audio-v2 (>20k hours of Hebrew audio) cut into 2–30 second speech segments with machine transcripts, ready for ASR fine-tuning. How it was built VAD — Silero VAD (ONNX) over each episode decoded to 16 kHz mono. Speech regions longer than 30 s are split at the quietest sufficiently-long pause inside the window, so cuts land in silence rather than mid-word. Regions shorter than 2 s are dropped. Transcription —… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/ivirits-audio-v2-30s.audioautomatic-speech-recognition1M<n<10M0 likes775 downloads2mo agoHugging Face18Reza2kn /nasle-mana-clean-chunked-30s Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk. Configuration Rows Columns Meaning labeled (train/) 9,886 audio, label Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.audioautomatic-speech-recognition1K<n<10K0 likes697 downloads27d agoHugging Face19Reza2kn /nasle-mana-clean-chunked-30s-avasanj Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here. Splits Split Rows Audio Columns labeled 4,981 41.41 hours audio, label to_transcribe 11,127 92.72 hours audio The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.audioautomatic-speech-recognition10K<n<100K1 likes676 downloads19d agoHugging Face20NathanRoll /global-news-radio-30s Global News Radio Dataset Multilingual news radio recordings from 51 languages across 42 countries. Recordings 51 Total audio 1500 min (25.0 h) Format MP3 16kHz mono 64kbps Parquet shards 11 Languages 51 Countries 42 Size 687 MB Languages Amharic, Arabic, Bashkir, Basque, Belarusian, Bengali, Brazilian Portuguese,Portugues Do Brasil,Português Brasil, Catalan, Croatian, Czech, Danish, Dutch, English, Estonian, Faroese, Finnish, Flemish… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-30s.audioautomatic-speech-recognitionn<1K0 likes637 downloads6mo agoHugging Face21WhissleAI /Meta_STT_ZH_AIShell3 Meta Speech Recognition Mandarin Dataset (AISHELL3) This dataset contains both metadata and audio files for Mandarin speech recognition samples from the AISHELL3 corpus. Dataset Statistics Splits and Sample Counts train: 60098 samples valid: 3163 samples test: 24772 samples Example Samples train { "audio_filepath": "/external4/datasets/Mandarin/AISHELL3/wavs_train/SSB00430356.wav", "text": "她以 ENTITY_PRODUCT 滴鸡精 END 调养身体。 AGE_14_25… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_ZH_AIShell3.audioautomatic-speech-recognition10K<n<100K0 likes443 downloads1y agoHugging Face22projecte-aina /parlament_parla_v3 Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.audioautomatic-speech-recognition100K<n<1M1 likes418 downloads2y agoHugging Face23novelwolde36 /WaxalNLP Waxal Datasets The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus. Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and language… See the full description on the dataset page: https://huggingface.co/datasets/novelwolde36/WaxalNLP.audioautomatic-speech-recognition1M<n<10M1 likes413 downloads7mo agoHugging Face24KeisukeMiyamoto /nhk-archive-audio-30sgated NHK Archives Audio 30s This is a Japanese speech corpus derived from NHK Archives Audio. Audio from public NHK Archives records was segmented into clips of up to 30 seconds using voice activity detection. The dataset contains 344,722 accepted clips, totaling 1,661.01 hours. Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text contains LLM-assisted corrections based on the transcript and available source title and description. This… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio-30s.audioautomatic-speech-recognition100K<n<1M0 likes371 downloads1d agoHugging Face25hanamizuki-ai /genshin-voice-v3.4-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.4-mandarin.audiotext-to-speech10K<n<100K11 likes344 downloads3y agoHugging Face26MightyStudent /Egyptian-ASR-MGB-3 Egyptian Arabic dialect automatic speech recognition Dataset Summary This dataset was collected, cleaned and adjusted for huggingface hub and ready to be used for whisper finetunning/training. From MGB-3 website: The MGB-3 is using 16 hours multi-genre data collected from different YouTube channels. The 16 hours have been manually transcribed. The chosen Arabic dialect for this year is Egyptian. Given that dialectal Arabic has no orthographic rules, each program has… See the full description on the dataset page: https://huggingface.co/datasets/MightyStudent/Egyptian-ASR-MGB-3.audioautomatic-speech-recognition1K<n<10K23 likes334 downloads2y agoHugging Face27urarik /thchs30audioautomatic-speech-recognition10K<n<100K2 likes333 downloads4mo agoHugging Face28suleiman2003 /W_hausa_v3 Cleaned Hausa Speech Dataset v3 A cleaned and processed Hausa speech dataset built from multiple open-source Hugging Face datasets. Dataset Description This dataset contains cleaned, normalized, and deduplicated Hausa speech audio with aligned transcriptions. All audio is: Sample rate: 16,000 Hz (mono) Format: FLAC (lossless, embedded in Parquet) Duration range: 1–30 seconds per clip Loudness normalized: -20 dBFS RMS VAD trimmed: Non-speech segments removed with… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/W_hausa_v3.audioautomatic-speech-recognition100K<n<1M0 likes329 downloads2mo agoHugging Face29pourmand1376 /asr-farsi-youtube-chunked-30-seconds How To Use from datasets import load_dataset train = load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='train+val') test =load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='test') +300 Hours ASR dataset generated from this kaggle dataset audioautomatic-speech-recognition10K<n<100K11 likes308 downloads3y agoHugging Face30MohamedRashad /MGB-3-Arabic Dataset Card for MGB-3 Arabic Speech Recognition Dataset Summary The MGB-3 Arabic dataset is a multi-genre collection of Egyptian Arabic speech extracted from YouTube videos, designed for speech recognition in challenging, real-world conditions. Unlike its predecessor MGB-2 which focused on broadcast TV news, MGB-3 emphasizes dialectal Arabic across diverse content types. The dataset contains approximately 16 hours of Egyptian Arabic speech from 80 YouTube videos… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MGB-3-Arabic.audioautomatic-speech-recognition1K<n<10K7 likes287 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.