datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
african_speech_cleanCommon-Voice-Speech-26.0-Persian-Clean
Persian Common Voice Clean Dataset
This dataset is a cleaned and prepared subset of the Persian (فارسی - fa) portion of Mozilla Common Voice Scripted Speech, based on cv-corpus-26.0-2026-06-12.
The cleaned release contains 34,134 audio clips, representing approximately 43.105 hours of speech, equal to 2,586.303 minutes. The clips are associated with approximately 34,134 validated Persian sentences and come from 3,791 speakers.
The original Persian Common Voice release contains… See the full description on the dataset page: https://huggingface.co/datasets/pymmdrza/Common-Voice-Speech-26.0-Persian-Clean.clean-speech-raw-sources
Clean Speech — Raw Source Archives (mirror)
Durable public mirrors of speech corpora whose original home is not Hugging Face
(external academic hosts disappear, move, or go offline). HF is used only as a
faster/durable mirror — the original source links are below and remain the canonical
home. Every archive is byte-for-byte unmodified from its origin and retains its original
licence and attribution.
42 corpora · 297 GB · 171+ files.
Scope note. This mirror began as… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/clean-speech-raw-sources.Clean_Common_Voice_Speech_24.0-TW
Cleaned Common Voice 24.0 - Chinese (Taiwan) Voice Seeds
資料集簡介 (Dataset Summary)
本資料集基於 Mozilla Common Voice Scripted Speech 24.0 - Chinese (Taiwan) 進行二次加工與清洗。
主要目的是萃取出高品質、無冗長空白、且長度適中的「聲色種子 (Voice Seeds)」,非常適合用於訓練或微調文字轉語音 (TTS)、語音複製 (Voice Cloning) 等生成式語音模型。
處理流程 (Data Processing Pipeline)
原始的 Common Voice 資料包含許多長短不一、可能帶有環境噪音或冗長靜音的音檔。本專案透過以下自動化流程進行清洗:
語音活動偵測 (VAD) 與去空白:
採用 silero-vad 模型進行精準的語音區段偵測 (Threshold: 0.5)。
自動剔除音檔前後與句間的冗長靜音,僅保留清晰的語音內容。
target_sr: int… See the full description on the dataset page: https://huggingface.co/datasets/OKHand/Clean_Common_Voice_Speech_24.0-TW.peoples-speech-asr-clean
peoples-speech-asr-clean
Filtered ASR dataset. Samples with <3 words, repetitive tokens, or chat token leaks removed.
twi-speech-text-multispeaker-cleanpeoples_speech-cleanThe People's Speech is a free-to-download 30,000-hour and growing supervised
conversational English speech recognition dataset licensed for academic and
commercial usage under CC-BY-SA (with a CC-BY subset).Egyptian-Speech-Clean-MGB3
🏛️ Dataset Card for MGB3-Egyptian-Clean
Dataset Summary
This dataset is a refined and enhanced version of the MGB-3 (Multi-Genre Broadcast) corpus, specifically focused on the Egyptian Arabic dialect. It has been meticulously preprocessed to be "TTS-ready" by combining advanced deep-learning denoising with custom linguistic text normalization.
🛠️ Preprocessing Pipeline
To ensure the highest quality for generative speech tasks (like VITS or MMS finetuning)… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/Egyptian-Speech-Clean-MGB3.speeches-congre-clean
Dataset Card for "speeches-congre-clean"
More Information needed
mixed-instruction-speech-multiturn-noise-cleanbambara-speech-kis-clean-split
Bambara Speech Dataset — Clean & Split
Dataset de reconnaissance vocale en bambara, nettoyé et splitté pour le fine-tuning de modèles ASR (ex: Whisper).
La source principale des données brutes est RobotsMali/bam-asr-early, auquel un remerciement chaleureux lui est attribué mais aussi à d'autres personnes référencées ci-dessous dans la section citation.
Statistiques
Total : 35 342 échantillons
Train : 24 738
Validation : 3 535
Test : 7 069
Durée moyenne : 3.23s… See the full description on the dataset page: https://huggingface.co/datasets/kalilouisangare/bambara-speech-kis-clean-split.twi-trigrams-speech-text-parallel-cleanspeech-sensor-fusion-clean
Speech Sensor Fusion Data Notes
Dataset summary
Preparation notes and schema examples for Speech tasks using Sensor Fusion data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/nicolasmarva/speech-sensor-fusion-clean.peoples_speech-clean-timestampedThe People's Speech is a free-to-download 30,000-hour and growing supervised
conversational English speech recognition dataset licensed for academic and
commercial usage under CC-BY-SA (with a CC-BY subset).new-moore-speech-clean
Moore Speech Proverbs: A Parallel Audio-Text Corpus for Mooré and French
The Moore Speech Proverbs dataset is a bilingual audio-text corpus of traditional proverbs in Mooré and French, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language.
[!NOTE]
⚠️ Access is gated. To request access, please read the… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-moore-speech-clean.speeches-congre-clean-names
Dataset Card for "speeches-congre-clean-names"
More Information needed
new-dioula-speech-clean
Dioula Speech Corpus: A Parallel Audio-Text Dataset for Dioula and French
The Dioula Speech Corpus is a bilingual audio-text corpus designed for research and academic purposes in low-resource speech and language processing. It is intended primarily to support the development of Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models for the Dioula language.
⚠️ Access is gated. To request access, please read the policy below.🛑 TLDR: For safety and traceability reasons… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-dioula-speech-clean.urdu_clean_speechclean_khmer_mpwt_speechThe initial implementation inspired by KrorngAI
clean_khmer_mpwt_speechcv10-uk-testset-clean-zipaGenerated by https://github.com/lingjzhu/zipa (zipa_large_crctc_500000_avg10.pth) and https://github.com/dmort27/epitran
peoples-speech-clean-Kanade-Annotatedspeech-dataset-public-cleanleaderboard-speech-cleanspeech-dataset-cleankhmer-mpwt-speech-cleanpeoples_speech-cleannoisy_clean_indic_speechclean-luvox-speechMixed_dataset_speech_sliced_clean
