CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01XRXRX /X-Voice-Dataset-Train X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling. Also the train set of X-Voice Model. Core Statistics Total Speech Duration: 420K hours 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.audiotext-to-speech10M<n<100M11 likes4.7k downloads5mo agoHugging Face02ivrit-ai /knesset-plenums-whisper-traininggated Dataset Card for ivrit.ai - Knesset Plenums Whisper Training This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset. This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less. Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription. The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.audiotext-to-speech100K<n<1M3 likes681 downloads10mo agoHugging Face03KitTzk /lao_stt_training_data Lao Speech-to-Text Training Data ຊຸດຂໍ້ມູນນີ້ຖືກຈັດກຽມຂຶ້ນມາເພື່ອໃຊ້ສຳລັບການເທຣນ ແລະ ປັບແຕ່ງ (Fine-tuning) ໂມເດວ Speech-to-Text (ເຊັ່ນ OpenAI Whisper) ສຳລັບພາສາລາວ. ໂຄງສ້າງຂອງຂໍ້ມູນ (Dataset Structure) Train set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ train/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ train.csv Validation set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ validation/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ validation.csv ຮູບແບບຂໍ້ມູນໃນໄຟລ໌ CSV: audio: ເສັ້ນທາງໄປຫາໄຟລ໌ສຽງ (e.g., train/audio25000.wav)… See the full description on the dataset page: https://huggingface.co/datasets/KitTzk/lao_stt_training_data.audioautomatic-speech-recognition1K<n<10K1 likes158 downloads3mo agoHugging Face04danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes144 downloads10mo agoHugging Face05ivrit-ai /crowd-recital-whisper-traininggated Dataset Card for ivrit.ai - Crowd Recital Dataset Details Dataset Description License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. - Full license: https://www.ivrit.ai/en/the-license/ - FAQs: https://www.ivrit.ai/en/license-faqs/ Dataset Structure Data Fields Each example in the dataset contains: audio: An audio column containing: bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.audiotext-to-speech1K<n<10K3 likes81 downloads10mo agoHugging Face06Tnaot /whisper-large-training Khmer Speech Dataset for Whisper Large V3 Combined Khmer speech datasets for fine-tuning Whisper models. Statistics Total: 52,255 samples (~47.3 hours) Train: 49,119 (94.0%) Test: 3,136 (6.0%) Duration: 1.0s - 13.3s (avg: 3.3s) Sources seanghay/khmer_mpwt_speech (×5) Samples: 10,290 (duplicated 5x from 2,058) Duration: ~9.6 hours seanghay/km-speech-corpus Samples: 14,943 Duration: ~10.3 hours google/fleurs Samples: 1,675… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/whisper-large-training.audioautomatic-speech-recognition10K<n<100K0 likes50 downloads11mo agoHugging Face07nlpctx /tts-training-dataset Human Reviewed Telugu-English TTS Dataset A manually reviewed multilingual TTS dataset created from publicly available educational and speech content. Dataset Splits & Distribution Metrics balanced_60min Split Total Segments: 120 Total Duration: 60.00 minutes Unique Speakers: 3 Distribution Breakdowns: Language Distribution: en-IN: 60 segments (30.00 minutes) te-IN: 60 segments (30.00 minutes) Style Distribution: analytical: 27… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tts-training-dataset.audiotext-to-speechn<1K0 likes49 downloads3mo agoHugging Face08ivrit-ai /crowd-recital-yi-whisper-traininggated Dataset Card for ivrit.ai - Crowd Recital - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~78h License: other Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.audiotext-to-speech10K<n<100K0 likes48 downloads10mo agoHugging Face09ivrit-ai /crowd-whatsapp-yi-whisper-traininggated Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~19h License: other Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.audiotext-to-speech1K<n<10K0 likes26 downloads10mo agoHugging Face10zionia /isizulu-asr-train isiZulu Speech Recognition Augmented Train Dataset Dataset Description This dataset contains augmented speech recordings and transcriptions for isiZulu, one of South Africa's official languages. The dataset has been optimized for use with OpenAI's Whisper ASR models. Dataset Statistics Number of samples: 690 Language: isiZulu (Zul) Audio format: WAV, 16kHz, mono, 16-bit Maximum duration: 30 seconds (truncated for Whisper compatibility) Transcription format:… See the full description on the dataset page: https://huggingface.co/datasets/zionia/isizulu-asr-train.audioautomatic-speech-recognitionn<1K0 likes23 downloads11mo agoHugging Face11norjas1 /trainSetaudioautomatic-speech-recognition1K<n<10K0 likes22 downloads2y agoHugging Face12zionia /isixhosa-asr-train isiXhosa Speech Recognition Augmented Dataset Dataset Description This dataset contains augmented speech recordings and transcriptions for isiXhosa, one of South Africa's official languages. The dataset has been optimized for use with OpenAI's Whisper ASR models. Dataset Statistics Number of samples: 680 Language: isiXhosa (Xho) Audio format: WAV, 16kHz, mono, 16-bit Maximum duration: 30 seconds (truncated for Whisper compatibility) Transcription format:… See the full description on the dataset page: https://huggingface.co/datasets/zionia/isixhosa-asr-train.audioautomatic-speech-recognitionn<1K0 likes15 downloads10mo agoHugging Face13SRP-base-model-training /kazakh_speech_corpus_2gated Kazakh_speech_dataset_2 This dataset contains Kazakh_speech_dataset_2 from ISSAI but in parquet format. Dataset info 645,860 Utterances 1194 Hours in total Sources in each split: test : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} train : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} validation : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts','podcasts'} Guides… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_corpus_2.audioautomatic-speech-recognition100K<n<1M2 likes14 downloads1y agoHugging Face14SRP-base-model-training /kazakh_speech_dataset_ksdgatedKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api. Dataset info: 813 Speakers with 500 samples for 4 speakers with 250 samples for 809 speakers Male/female 555 Hours Guides Load data 1 Replace the export HF_HOME with your HF_HOME path from datasets import load_dataset # export HF_HOME="/data/vladimir_albrekht/hf_cache" ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.audioautomatic-speech-recognition100K<n<1M2 likes14 downloads1y agoHugging Face15wasertech /TrainingSpeechTrainingSpeech is an initiative to provide open and freely reusable dataset of voices for speech-to-text models training on non-english languages using already available data (such as audio-books). Right now, data are extracted exclusively from audio-books and in French language. Let me know if you are intersted to contribute by creating an issue. Tooling TrainingSpeech comes with a CLI that automate and simplify: transcript extraction forced-alignment (using aeneas)… See the full description on the dataset page: https://huggingface.co/datasets/wasertech/TrainingSpeech.audioautomatic-speech-recognition100K<n<1M2 likes10 downloads1y agoHugging Face16WhissleAI /Meta_STT_MADASR2.0_train_lggated Dataset Card for Dataset Name https://sites.google.com/view/respinasrchallenge2025/home?authuser=0 This is the trainig dataset provided for track-3 and track-4 of this task. We enhance the dataset with entity tagging, emotion, age, gender and intent. WhissleAI participated in this challenge. This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_MADASR2.0_train_lg.audioaudio-classification100K<n<1M0 likes6 downloads4mo agoHugging Face17MOH749 /bocalantics-wol-train Bocalantics Wolof eval set The held-out Wolof test rows from MOH749/Bocalantics-2.0, with audio attached, so a candidate model can be scored without re-materialising anything. 43,034 rows, 65.33 hours. Why it exists "Beats unadapted Whisper" is not a result for Wolof. Whisper has seen 2 of the parent corpus's 26 languages, so beating it is arithmetic rather than evidence. The bar is the best published model for the language -… See the full description on the dataset page: https://huggingface.co/datasets/MOH749/bocalantics-wol-train.audioautomatic-speech-recognition10K<n<100K0 likes6 downloads19d agoHugging Face18gallip0li /medimind-r11-traingated MediMind R11 — ASR training data Unified manifest + packed audio for fine-tuning Whisper-large-v3 on Norwegian clinical and conversational speech. Training manifest: r11_manifest.jsonl — 11,022 packs Held-out eval set: r11_heldout_eval.jsonl — 291 packs (NEVER train on these) ~see manifest audit packs total 11 sources: lege_*, podcasts (motiv/podk/stet), nb_samtale, nb_tale_m3, tts_drugs Schema See r11_manifest.jsonl (one JSON object per line) and… See the full description on the dataset page: https://huggingface.co/datasets/gallip0li/medimind-r11-train.audioautomatic-speech-recognitionn<1K0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.