CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anuj-inavlabs /Thinkspark-v2-270m-training-data ThinkSpark-v2-350M — training data Full-duplex floor-controller (Section 8) training corpus: playable audio + text, paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain, gender, prosody, agent text) and Soniox character-level timestamps. Dataset Viewer Default split is parquet with a real Audio feature — a player renders inline next to the text in the Hub UI: column type description audio Audio playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.text-to-speech1K<n<10K0 likes6.9k downloads22d agoHugging Face02ivrit-ai /knesset-plenums-whisper-traininggated Dataset Card for ivrit.ai - Knesset Plenums Whisper Training This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset. This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less. Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription. The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.audiotext-to-speech100K<n<1M3 likes559 downloads10mo agoHugging Face03danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes142 downloads10mo agoHugging Face04ivrit-ai /crowd-recital-whisper-traininggated Dataset Card for ivrit.ai - Crowd Recital Dataset Details Dataset Description License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. - Full license: https://www.ivrit.ai/en/the-license/ - FAQs: https://www.ivrit.ai/en/license-faqs/ Dataset Structure Data Fields Each example in the dataset contains: audio: An audio column containing: bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.audiotext-to-speech1K<n<10K3 likes81 downloads10mo agoHugging Face05KitTzk /lao_stt_training_data Lao Speech-to-Text Training Data ຊຸດຂໍ້ມູນນີ້ຖືກຈັດກຽມຂຶ້ນມາເພື່ອໃຊ້ສຳລັບການເທຣນ ແລະ ປັບແຕ່ງ (Fine-tuning) ໂມເດວ Speech-to-Text (ເຊັ່ນ OpenAI Whisper) ສຳລັບພາສາລາວ. ໂຄງສ້າງຂອງຂໍ້ມູນ (Dataset Structure) Train set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ train/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ train.csv Validation set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ validation/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ validation.csv ຮູບແບບຂໍ້ມູນໃນໄຟລ໌ CSV: audio: ເສັ້ນທາງໄປຫາໄຟລ໌ສຽງ (e.g., train/audio25000.wav)… See the full description on the dataset page: https://huggingface.co/datasets/KitTzk/lao_stt_training_data.audioautomatic-speech-recognition1K<n<10K1 likes65 downloads3mo agoHugging Face06nlpctx /tts-training-dataset Human Reviewed Telugu-English TTS Dataset A manually reviewed multilingual TTS dataset created from publicly available educational and speech content. Dataset Splits & Distribution Metrics balanced_60min Split Total Segments: 120 Total Duration: 60.00 minutes Unique Speakers: 3 Distribution Breakdowns: Language Distribution: en-IN: 60 segments (30.00 minutes) te-IN: 60 segments (30.00 minutes) Style Distribution: analytical: 27… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tts-training-dataset.audiotext-to-speechn<1K0 likes47 downloads3mo agoHugging Face07Tnaot /whisper-large-training Khmer Speech Dataset for Whisper Large V3 Combined Khmer speech datasets for fine-tuning Whisper models. Statistics Total: 52,255 samples (~47.3 hours) Train: 49,119 (94.0%) Test: 3,136 (6.0%) Duration: 1.0s - 13.3s (avg: 3.3s) Sources seanghay/khmer_mpwt_speech (×5) Samples: 10,290 (duplicated 5x from 2,058) Duration: ~9.6 hours seanghay/km-speech-corpus Samples: 14,943 Duration: ~10.3 hours google/fleurs Samples: 1,675… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/whisper-large-training.audioautomatic-speech-recognition10K<n<100K0 likes45 downloads11mo agoHugging Face08DigiGreen /KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers. Process of collection of data: Selected users were given the option of doing a task and getting paid for it. The users were supposed to record the sentence as it appeared on the screen. The audio file thus obtained was validated matched with the sentences to fine tune the model. Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/KikuyuASR_trainingdataset.automatic-speech-recognition10K<n<100K0 likes32 downloads2y agoHugging Face09ivrit-ai /crowd-recital-yi-whisper-traininggated Dataset Card for ivrit.ai - Crowd Recital - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~78h License: other Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.audiotext-to-speech10K<n<100K0 likes22 downloads10mo agoHugging Face10ivrit-ai /crowd-whatsapp-yi-whisper-traininggated Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~19h License: other Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.audiotext-to-speech1K<n<10K0 likes20 downloads10mo agoHugging Face11jo-05 /KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers. Process of collection of data: Selected users were given the option of doing a task and getting paid for it. The users were supposed to record the sentence as it appeared on the screen. The audio file thus obtained was validated matched with the sentences to fine tune the model. Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/jo-05/KikuyuASR_trainingdataset.automatic-speech-recognition10K<n<100K0 likes13 downloads7mo agoHugging Face12SRP-base-model-training /kazakh_speech_corpus_2gated Kazakh_speech_dataset_2 This dataset contains Kazakh_speech_dataset_2 from ISSAI but in parquet format. Dataset info 645,860 Utterances 1194 Hours in total Sources in each split: test : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} train : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} validation : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts','podcasts'} Guides… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_corpus_2.audioautomatic-speech-recognition100K<n<1M2 likes12 downloads1y agoHugging Face13SRP-base-model-training /kazakh_speech_dataset_ksdgatedKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api. Dataset info: 813 Speakers with 500 samples for 4 speakers with 250 samples for 809 speakers Male/female 555 Hours Guides Load data 1 Replace the export HF_HOME with your HF_HOME path from datasets import load_dataset # export HF_HOME="/data/vladimir_albrekht/hf_cache" ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.audioautomatic-speech-recognition100K<n<1M2 likes12 downloads1y agoHugging Face14wasertech /TrainingSpeechTrainingSpeech is an initiative to provide open and freely reusable dataset of voices for speech-to-text models training on non-english languages using already available data (such as audio-books). Right now, data are extracted exclusively from audio-books and in French language. Let me know if you are intersted to contribute by creating an issue. Tooling TrainingSpeech comes with a CLI that automate and simplify: transcript extraction forced-alignment (using aeneas)… See the full description on the dataset page: https://huggingface.co/datasets/wasertech/TrainingSpeech.audioautomatic-speech-recognition100K<n<1M2 likes10 downloads1y agoHugging Face15CGIAR /KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers. Process of collection of data: Selected users were given the option of doing a task and getting paid for it. The users were supposed to record the sentence as it appeared on the screen. The audio file thus obtained was validated matched with the sentences to fine tune the model. Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/KikuyuASR_trainingdataset.automatic-speech-recognition10K<n<100K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.