CoolFace
16 results

clean-speech

amanuelbyte /african_speech_cleanaudio1M<n<10M1 likes236 downloads5mo agoHugging Facepymmdrza /Common-Voice-Speech-26.0-Persian-Clean Persian Common Voice Clean Dataset This dataset is a cleaned and prepared subset of the Persian (فارسی - fa) portion of Mozilla Common Voice Scripted Speech, based on cv-corpus-26.0-2026-06-12. The cleaned release contains 34,134 audio clips, representing approximately 43.105 hours of speech, equal to 2,586.303 minutes. The clips are associated with approximately 34,134 validated Persian sentences and come from 3,791 speakers. The original Persian Common Voice release contains… See the full description on the dataset page: https://huggingface.co/datasets/pymmdrza/Common-Voice-Speech-26.0-Persian-Clean.audiotext-to-speech10K<n<100K2 likes233 downloads1mo agoHugging FaceScicom-intl /clean-speech-raw-sources Clean Speech — Raw Source Archives (mirror) Durable public mirrors of speech corpora whose original home is not Hugging Face (external academic hosts disappear, move, or go offline). HF is used only as a faster/durable mirror — the original source links are below and remain the canonical home. Every archive is byte-for-byte unmodified from its origin and retains its original licence and attribution. 42 corpora · 297 GB · 171+ files. Scope note. This mirror began as… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/clean-speech-raw-sources.0 likes205 downloads1mo agoHugging FaceOKHand /Clean_Common_Voice_Speech_24.0-TW Cleaned Common Voice 24.0 - Chinese (Taiwan) Voice Seeds 資料集簡介 (Dataset Summary) 本資料集基於 Mozilla Common Voice Scripted Speech 24.0 - Chinese (Taiwan) 進行二次加工與清洗。 主要目的是萃取出高品質、無冗長空白、且長度適中的「聲色種子 (Voice Seeds)」,非常適合用於訓練或微調文字轉語音 (TTS)、語音複製 (Voice Cloning) 等生成式語音模型。 處理流程 (Data Processing Pipeline) 原始的 Common Voice 資料包含許多長短不一、可能帶有環境噪音或冗長靜音的音檔。本專案透過以下自動化流程進行清洗: 語音活動偵測 (VAD) 與去空白: 採用 silero-vad 模型進行精準的語音區段偵測 (Threshold: 0.5)。 自動剔除音檔前後與句間的冗長靜音,僅保留清晰的語音內容。 target_sr: int… See the full description on the dataset page: https://huggingface.co/datasets/OKHand/Clean_Common_Voice_Speech_24.0-TW.audio10K<n<100K2 likes106 downloads7mo agoHugging FaceOpenSpeechHub /peoples-speech-asr-clean peoples-speech-asr-clean Filtered ASR dataset. Samples with <3 words, repetitive tokens, or chat token leaks removed. audio100K<n<1M1 likes102 downloads6mo agoHugging Faceghanaopenai /twi-speech-text-multispeaker-cleanaudio1K<n<10K0 likes53 downloads10mo agoHugging Face