CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amanuelbyte /african_speech_cleanaudio1M<n<10M1 likes236 downloads5mo agoHugging Face02pymmdrza /Common-Voice-Speech-26.0-Persian-Clean Persian Common Voice Clean Dataset This dataset is a cleaned and prepared subset of the Persian (فارسی - fa) portion of Mozilla Common Voice Scripted Speech, based on cv-corpus-26.0-2026-06-12. The cleaned release contains 34,134 audio clips, representing approximately 43.105 hours of speech, equal to 2,586.303 minutes. The clips are associated with approximately 34,134 validated Persian sentences and come from 3,791 speakers. The original Persian Common Voice release contains… See the full description on the dataset page: https://huggingface.co/datasets/pymmdrza/Common-Voice-Speech-26.0-Persian-Clean.audiotext-to-speech10K<n<100K2 likes233 downloads1mo agoHugging Face03Scicom-intl /clean-speech-raw-sources Clean Speech — Raw Source Archives (mirror) Durable public mirrors of speech corpora whose original home is not Hugging Face (external academic hosts disappear, move, or go offline). HF is used only as a faster/durable mirror — the original source links are below and remain the canonical home. Every archive is byte-for-byte unmodified from its origin and retains its original licence and attribution. 42 corpora · 297 GB · 171+ files. Scope note. This mirror began as… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/clean-speech-raw-sources.0 likes205 downloads1mo agoHugging Face04OKHand /Clean_Common_Voice_Speech_24.0-TW Cleaned Common Voice 24.0 - Chinese (Taiwan) Voice Seeds 資料集簡介 (Dataset Summary) 本資料集基於 Mozilla Common Voice Scripted Speech 24.0 - Chinese (Taiwan) 進行二次加工與清洗。 主要目的是萃取出高品質、無冗長空白、且長度適中的「聲色種子 (Voice Seeds)」,非常適合用於訓練或微調文字轉語音 (TTS)、語音複製 (Voice Cloning) 等生成式語音模型。 處理流程 (Data Processing Pipeline) 原始的 Common Voice 資料包含許多長短不一、可能帶有環境噪音或冗長靜音的音檔。本專案透過以下自動化流程進行清洗: 語音活動偵測 (VAD) 與去空白: 採用 silero-vad 模型進行精準的語音區段偵測 (Threshold: 0.5)。 自動剔除音檔前後與句間的冗長靜音,僅保留清晰的語音內容。 target_sr: int… See the full description on the dataset page: https://huggingface.co/datasets/OKHand/Clean_Common_Voice_Speech_24.0-TW.audio10K<n<100K2 likes106 downloads7mo agoHugging Face05OpenSpeechHub /peoples-speech-asr-clean peoples-speech-asr-clean Filtered ASR dataset. Samples with <3 words, repetitive tokens, or chat token leaks removed. audio100K<n<1M1 likes102 downloads6mo agoHugging Face06ghanaopenai /twi-speech-text-multispeaker-cleanaudio1K<n<10K0 likes53 downloads10mo agoHugging Face07distil-whisper /peoples_speech-cleanThe People's Speech is a free-to-download 30,000-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA (with a CC-BY subset).automatic-speech-recognition1 likes51 downloads3y agoHugging Face08MohamedGomaa30 /Egyptian-Speech-Clean-MGB3 🏛️ Dataset Card for MGB3-Egyptian-Clean Dataset Summary This dataset is a refined and enhanced version of the MGB-3 (Multi-Genre Broadcast) corpus, specifically focused on the Egyptian Arabic dialect. It has been meticulously preprocessed to be "TTS-ready" by combining advanced deep-learning denoising with custom linguistic text normalization. 🛠️ Preprocessing Pipeline To ensure the highest quality for generative speech tasks (like VITS or MMS finetuning)… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/Egyptian-Speech-Clean-MGB3.audio1K<n<10K0 likes43 downloads8mo agoHugging Face09Sleoruiz /speeches-congre-clean Dataset Card for "speeches-congre-clean" More Information needed text10K<n<100K0 likes34 downloads4y agoHugging Face10jan-hq /mixed-instruction-speech-multiturn-noise-cleantabular100K<n<1M0 likes33 downloads2y agoHugging Face11kalilouisangare /bambara-speech-kis-clean-split Bambara Speech Dataset — Clean & Split Dataset de reconnaissance vocale en bambara, nettoyé et splitté pour le fine-tuning de modèles ASR (ex: Whisper). La source principale des données brutes est RobotsMali/bam-asr-early, auquel un remerciement chaleureux lui est attribué mais aussi à d'autres personnes référencées ci-dessous dans la section citation. Statistiques Total : 35 342 échantillons Train : 24 738 Validation : 3 535 Test : 7 069 Durée moyenne : 3.23s… See the full description on the dataset page: https://huggingface.co/datasets/kalilouisangare/bambara-speech-kis-clean-split.audio10K<n<100K2 likes33 downloads7mo agoHugging Face12michsethowusu /twi-trigrams-speech-text-parallel-cleanaudio10K<n<100K0 likes23 downloads10mo agoHugging Face13nicolasmarva /speech-sensor-fusion-clean Speech Sensor Fusion Data Notes Dataset summary Preparation notes and schema examples for Speech tasks using Sensor Fusion data. Full source material is intentionally not bundled, so provenance and licensing remain explicit. Included material preprocess.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable records for checking the schema. README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/nicolasmarva/speech-sensor-fusion-clean.0 likes23 downloads1mo agoHugging Face14distil-whisper /peoples_speech-clean-timestampedThe People's Speech is a free-to-download 30,000-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA (with a CC-BY subset).automatic-speech-recognition0 likes18 downloads3y agoHugging Face15goaicorp /new-moore-speech-cleangated Moore Speech Proverbs: A Parallel Audio-Text Corpus for Mooré and French The Moore Speech Proverbs dataset is a bilingual audio-text corpus of traditional proverbs in Mooré and French, designed for research and academic purposes in low-resource speech and language processing. It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language. [!NOTE] ⚠️ Access is gated. To request access, please read the… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-moore-speech-clean.audiotext-to-speech1K<n<10K1 likes18 downloads5mo agoHugging Face16Sleoruiz /speeches-congre-clean-names Dataset Card for "speeches-congre-clean-names" More Information needed text10K<n<100K0 likes15 downloads4y agoHugging Face17goaicorp /new-dioula-speech-cleangated Dioula Speech Corpus: A Parallel Audio-Text Dataset for Dioula and French The Dioula Speech Corpus is a bilingual audio-text corpus designed for research and academic purposes in low-resource speech and language processing. It is intended primarily to support the development of Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models for the Dioula language. ⚠️ Access is gated. To request access, please read the policy below.🛑 TLDR: For safety and traceability reasons… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-dioula-speech-clean.audiotext-to-speech10K<n<100K1 likes15 downloads5mo agoHugging Face18humair025 /urdu_clean_speechaudio10K<n<100K0 likes12 downloads1y agoHugging Face19Sophy15-St /clean_khmer_mpwt_speechThe initial implementation inspired by KrorngAI audion<1K0 likes10 downloads6mo agoHugging Face20Kimang18 /clean_khmer_mpwt_speechaudion<1K0 likes7 downloads1y agoHugging Face21speech-uk /cv10-uk-testset-clean-zipaGenerated by https://github.com/lingjzhu/zipa (zipa_large_crctc_500000_avg10.pth) and https://github.com/dmort27/epitran text1K<n<10K0 likes7 downloads11mo agoHugging Face22humair025 /peoples-speech-clean-Kanade-Annotatedtabular1K<n<10K0 likes7 downloads8mo agoHugging Face23burkimbia /speech-dataset-public-cleangatedaudio10K<n<100K0 likes7 downloads5mo agoHugging Face24burkimbia /leaderboard-speech-cleangatedaudion<1K0 likes6 downloads2mo agoHugging Face25burkimbia /speech-dataset-cleangatedaudio10K<n<100K0 likes5 downloads5mo agoHugging Face26KrorngAI /khmer-mpwt-speech-cleantext1K<n<10K0 likes4 downloads1y agoHugging Face27mitchelldehaven /peoples_speech-clean0 likes2 downloads2y agoHugging Face28rudrathakar18 /noisy_clean_indic_speech0 likes2 downloads1y agoHugging Face29luvox-ai /clean-luvox-speechgatedaudion<1K0 likes1 downloads7mo agoHugging Face30Artaudio /Mixed_dataset_speech_sliced_cleangatedtabular100K<n<1M0 likes1 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.