CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Maisum-Abbas-123 /Urdu-Finetuning-Data-VibeVoice-Largeaudio10K<n<100K0 likes179 downloads8mo agoHugging Face02laion /emotional-roleplay-finetuning-dataset Artificial Voice Roleplay Dataset 67,491 fully-synthetic speech clips (~184 hours) pairing expressive role-play / character voice-direction captions with generated audio, across German, English, Spanish, and French (German-dominant). Rich in exaggerated fantasy/creature voices (orc, goblin, troll, ogre, zombie, dragon, demon, witch, banshee, imp, fairy, gnome, robot, murloc, harpy, skeleton, ghost, vampire …) and high-arousal emotional delivery (rage, fear, grief, menace). Every… See the full description on the dataset page: https://huggingface.co/datasets/laion/emotional-roleplay-finetuning-dataset.audiotext-to-speech10K<n<100K5 likes142 downloads2mo agoHugging Face03baptistefrancois1 /s2s-fr-finetuning s2s-fr-finetuning Corpus FR pour le finetuning speech-to-speech (Liquid-Audio / LFM2-Audio), construit par une pipeline de prétraitement : VAD, ASR + alignement mot, segmentation aux frontières de mots, filtrage qualité perceptuelle, normalisation de texte, déduplication. Utilisation from datasets import load_dataset ds = load_dataset("baptistefrancois1/s2s-fr-finetuning", "common_voice_fr") Un config HF par source d'origine : common_voice_fr, emilia_yodas_fr… See the full description on the dataset page: https://huggingface.co/datasets/baptistefrancois1/s2s-fr-finetuning.audio100K<n<1M0 likes134 downloads1mo agoHugging Face04nickoo004 /FeruzaSpeech_to_fine_tuning FeruzaSpeech_to_fine_tuning A speech corpus of ⏱️ ~59.1 total hours of Uzbek audio paired with Latin‑script transcripts, intended for fine‑tuning ASR / speech‑to‑text models. Dataset Details Dataset Description This dataset contains recordings of native Uzbek speakers reading a mix of classical literature excerpts and school‑level writing prompts: 001: Choliqushi (a novel by Rashod Nuri Guntekin, trans. by Mirzakalon Ismoiliy; first pub. Sept 1900). 002:… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/FeruzaSpeech_to_fine_tuning.audioautomatic-speech-recognition10K<n<100K1 likes130 downloads1y agoHugging Face05demegire /personaplex-finetuning-pharma-data-sample PersonaPlex Finetuning — Pharma Data Sample A 10-example slice of the synthetic patient-support / medication adherence dataset used to train demegire/personaplex-finetune-pharma. The on-disk layout below is exactly what the trainer in emotion-machine-org/personaplex-finetune consumes — use this as a template when building your own. Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at sample scale). Layout . ├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.audiotext-to-speechn<1K0 likes68 downloads5mo agoHugging Face06MAdel121 /Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning Curated Arabic Speech Dataset for Seasme (from MCV17) Dataset Description This dataset is a curated and preprocessed version of the Arabic (ar) subset from Mozilla Common Voice (MCV) 17.0. It has been specifically prepared for fine-tuning conversational speech models, with a primary focus on the Seasme-CSM model architecture. The dataset consists of audio clips in WAV format (24kHz, mono) and their corresponding transcripts, along with integer speaker IDs. The original… See the full description on the dataset page: https://huggingface.co/datasets/MAdel121/Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning.audio10K<n<100K1 likes66 downloads1y agoHugging Face07OpenWhistleNeurIPS26 /OpenWhistle-Classification-Finetuning OpenWhistle Classification Finetuning Dataset OpenWhistleNeurIPS26/OpenWhistle-Classification-Finetuning is the public classification finetuning dataset used for dolphin whistle identity classification. It contains short whistle clips, whistle-level metadata, fundamental-frequency tracks, rendered F0 spectrograms, and integer class labels. The main reviewer-facing subset is the balanced balanced config. It contains six classes: NSW_1 (label=0) SW_Luna (label=1) SW_Nana (label=2)… See the full description on the dataset page: https://huggingface.co/datasets/OpenWhistleNeurIPS26/OpenWhistle-Classification-Finetuning.audioaudio-classification10K<n<100K0 likes60 downloads5mo agoHugging Face08AhmedRezik /fineTuningDataaudio10K<n<100K0 likes34 downloads3y agoHugging Face09Sk1382 /Indian_Englsih_SSML_dataset_for_orpheus_fine_tuningaudion<1K0 likes32 downloads1y agoHugging Face10shane062 /FYP_Fine_Tuning Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/shane062/FYP_Fine_Tuning.audioautomatic-speech-recognitionn<1K1 likes29 downloads2y agoHugging Face11tonypeng /whisper-finetuningaudio1K<n<10K0 likes29 downloads2y agoHugging Face12Dev523 /tts-finetuningaudion<1K0 likes19 downloads1y agoHugging Face13B0808 /MDbA_FineTuningaudion<1K0 likes16 downloads3y agoHugging Face14dngngnguyencode /TTS_finetuning_datasetaudio10K<n<100K0 likes16 downloads2y agoHugging Face15thomaslu /articulationGAN_finetuning_data Dataset Card for "articulationGAN_finetuning_data" More Information needed audion<1K0 likes15 downloads3y agoHugging Face16classen3 /whisper-finetuning-for-aseeaudion<1K0 likes15 downloads3y agoHugging Face17B0808 /MDbA-FineTuningV2audion<1K0 likes12 downloads3y agoHugging Face18Kady-x /nova_finetuningaudion<1K0 likes12 downloads1y agoHugging Face19roshbeed /ai-residency-whisper-fine-tuning-dataaudion<1K0 likes11 downloads1mo agoHugging Face20YongJaeLee /Whisper_FineTuning_Koaudio10K<n<100K0 likes10 downloads1y agoHugging Face21YongJaeLee /Whisper_FineTuning_Suaudio10K<n<100K0 likes10 downloads1y agoHugging Face22deboleen6 /whisper_finetuningaudion<1K0 likes9 downloads2y agoHugging Face23OleksandrAbashkin /fine-tuning-Jewish Dataset Card for "fine-tuning-Jewish" More Information needed audion<1K0 likes9 downloads2y agoHugging Face24wahyuachmad /Zeta-Voice-ID-Vits-Finetuningaudion<1K0 likes9 downloads2y agoHugging Face25sravan-gorugantu /techolution_noun_finetuning_v1audion<1K0 likes9 downloads2y agoHugging Face26tonypeng /whisper-finetuning-testaudion<1K0 likes9 downloads2y agoHugging Face27anchaeyeon /whisper_finetuningaudio1K<n<10K0 likes9 downloads1y agoHugging Face28Shubham079 /whisper-finetuning-audio-filesaudio10K<n<100K0 likes8 downloads1y agoHugging Face29hostbot77 /Speech_to_fine_tuning FeruzaSpeech_to_fine_tuning A speech corpus of ⏱️ ~59.1 total hours of Uzbek audio paired with Latin‑script transcripts, intended for fine‑tuning ASR / speech‑to‑text models. Dataset Details Dataset Description This dataset contains recordings of native Uzbek speakers reading a mix of classical literature excerpts and school‑level writing prompts: 001: Choliqushi (a novel by Rashod Nuri Guntekin, trans. by Mirzakalon Ismoiliy; first pub. Sept 1900). 002:… See the full description on the dataset page: https://huggingface.co/datasets/hostbot77/Speech_to_fine_tuning.audioautomatic-speech-recognition10K<n<100K0 likes8 downloads8mo agoHugging Face30LasseRogers2111 /stt_lowrank_finetuningaudion<1K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.