CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes609 downloads17d agoHugging Face02Sheeba2026 /bharatvani-hindi-speech-corpusgated BharatVani Hindi Speech Corpus (150-Hour Studio Dataset) Proprietary Speech Asset • TheCreatorOS • BharatVani AI 1. Overview The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi. Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.audiotext-to-speech100K<n<1M1 likes284 downloads5d agoHugging Face03Archit00 /aijockey-public-corpusaudion<1K0 likes197 downloads4mo agoHugging Face04FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes175 downloads6mo agoHugging Face05Nam-toon-studio /Gurbani-MahanKosh-Frontier-Corpus ੴ Gurbani & Bhai Kahn Singh Nabha Mahan Kosh Frontier Corpus ☬ ਗੁਰਬਾਣੀ ਅਤੇ ਭਾਈ ਕਾਹਨ ਸਿੰਘ ਨਾਭਾ 'ਮਹਾਨ ਕੋਸ਼' ਪ੍ਰਮਾਣਿਕ ਡਾਟਾਸੈੱਟ 👨‍💻 Project Lead & Architecture Curator & Developer: Gurpreet Singh Dhillon (Nam-toon Studio) GitHub Profile: github.com/gurpreetsingh5523-source Flagship Project: AMRIT Research OS (Autonomous Medical AI) 📖 Dataset Overview An authoritative lexical dataset compiling authentic definitions… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Gurbani-MahanKosh-Frontier-Corpus.audiotext-generationn<1K0 likes80 downloads3d agoHugging Face06vocence /vocence_eval_corpus Dataset Card for vocence_eval_corpus Dataset Summary Everything produced by evaluating 7 PromptTTS systems (gemini, voxcpm, qwen3, maya1, openai, parler-tts, elevenlabs) on the 200-item balanced subset of vocence_corpus: the generated audio, every per-clip judge output, and the aggregated leaderboards/charts. For the prompt corpus alone (no eval data), see vocence_corpus. Scoring library: vocencebench (source) Benchmark code + full methodology write-up:… See the full description on the dataset page: https://huggingface.co/datasets/vocence/vocence_eval_corpus.audiotext-to-speechn<1K0 likes68 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.