CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huggingface-course /supervised-finetuning_quiz_student_responsestextn<1K4 likes1.3k downloads2h agoHugging Face02KaLM-Embedding /KaLM-embedding-finetuning-dataThe pretraining dataset is available at this link: HIT-TMG/KaLM-embedding-pretrain-data. Languages English, Chinese, Multilingual Dataset Structure Each in datasets is in the following format: query, string, one query per sample pos, list[string], usually containing one positive example neg, list[string], usually containing seven negative examples Dataset Summary All these datasets have been preprocessed and can be used for finetuning your embedding models.… See the full description on the dataset page: https://huggingface.co/datasets/KaLM-Embedding/KaLM-embedding-finetuning-data.textfeature-extraction1M<n<10M32 likes1k downloads10mo agoHugging Face03science-of-finetuning /fineweb-1m-sampletabular1M<n<10M1 likes581 downloads2y agoHugging Face04cpratikaki /RSVQA-HR_qwen_finetuningimage100K<n<1M1 likes541 downloads2y agoHugging Face05zihaojing /MuMo-Finetuning MuMo Finetuning Dataset This repository contains the finetuning datasets used in the paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning. Paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning Project Page: NeurIPS 2025 Poster Code: GitHub Repository Hub (this dataset): https://huggingface.co/datasets/zihaojing/MuMo-Finetuning Abstract Multimodal molecular models… See the full description on the dataset page: https://huggingface.co/datasets/zihaojing/MuMo-Finetuning.tabulargraph-ml100K<n<1M0 likes490 downloads11mo agoHugging Face06Makaareeem /publikasi-rag-finetuning-datasettext10K<n<100K1 likes407 downloads41m agoHugging Face07hellotayssir /FinQA_TAT-QA_financial_finetuning_dataset Dataset Summary This dataset provides a unified, flattened context / question / answer format for question answering over financial documents that combine tabular and textual data. It is built to support training and evaluating models on numerical and discrete reasoning tasks in the finance domain, drawing on the structure and style of established finance-QA benchmarks such as TAT-QA and FinQA. Each example pairs a passage of financial context (derived from a table and/or… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_TAT-QA_financial_finetuning_dataset.text10K<n<100K1 likes359 downloads2mo agoHugging Face08fine2006 /processed_dataset_whisper_finetuning10K<n<100K0 likes329 downloads1y agoHugging Face09YongJaeLee /Whisper_FineTuning_Su_preprocessing10K<n<100K0 likes299 downloads1y agoHugging Face10JustANormalTinkerer /hayai-finetuning-dataset-with-koreanimage10K<n<100K1 likes208 downloads11d agoHugging Face11Maisum-Abbas-123 /Urdu-Finetuning-Data-VibeVoice-Largeaudio10K<n<100K0 likes180 downloads8mo agoHugging Face12laion /emotional-roleplay-finetuning-dataset Artificial Voice Roleplay Dataset 67,491 fully-synthetic speech clips (~184 hours) pairing expressive role-play / character voice-direction captions with generated audio, across German, English, Spanish, and French (German-dominant). Rich in exaggerated fantasy/creature voices (orc, goblin, troll, ogre, zombie, dragon, demon, witch, banshee, imp, fairy, gnome, robot, murloc, harpy, skeleton, ghost, vampire …) and high-arousal emotional delivery (rage, fear, grief, menace). Every… See the full description on the dataset page: https://huggingface.co/datasets/laion/emotional-roleplay-finetuning-dataset.audiotext-to-speech10K<n<100K5 likes162 downloads2mo agoHugging Face13dgonier /Yusuf-OpenCaselist-finetuningtext1M<n<10M0 likes160 downloads2y agoHugging Face14fine2006 /unprocessed_dataset_whisper_finetuningtabular10K<n<100K0 likes160 downloads1y agoHugging Face15AnhMinhLe /refunc_fc_finetuningtext100K<n<1M0 likes157 downloads1y agoHugging Face16shawhin /tool-use-finetuningDataset for fine-tuning gemma-3-1b-it for function calling. The code and other resources for this project are linked below. Resources: YouTube Video Blog Post GitHub Repo Fine-tuned Model | Original Model Citation If you find this dataset helpful, please cite: @dataset{talebi2025, author = {Shaw Talebi}, title = {tool-use-finetuning}, year = {2025}, publisher = {Hugging Face}, howpublished =… See the full description on the dataset page: https://huggingface.co/datasets/shawhin/tool-use-finetuning.textn<1K25 likes155 downloads1y agoHugging Face17YongJaeLee /Whisper_FineTuning_Ko_preprocessing10K<n<100K0 likes152 downloads1y agoHugging Face18andresnowak /Instruction-finetuning-mixture-mnlpDataset created using the Tulu3-sft-mixture From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed Also the datasets for alignment and jailbreaking were removed text1M<n<10M0 likes151 downloads1y agoHugging Face19science-of-finetuning /lmsys-chat-1m-chat-formattedtext1M<n<10M0 likes146 downloads1y agoHugging Face20KaLM-Embedding /KaLM-embedding-finetuning-data-spanish KaLM-embedding-finetuning-data-spanish Spanish finetuning data for embedding models, adapted from the upstream dataset card of KaLM-Embedding/KaLM-embedding-finetuning-data. This directory contains a local Spanish version of the KaLM embedding finetuning corpus. It keeps the same training-oriented triplet/list structure as the upstream release and is organized as multiple parquet-backed subsets that can be loaded independently or combined for large-scale embedding training.… See the full description on the dataset page: https://huggingface.co/datasets/KaLM-Embedding/KaLM-embedding-finetuning-data-spanish.text1M<n<10M1 likes146 downloads5mo agoHugging Face21Mkid95 /sft_finetuning_dataset_tokenizedtext100K<n<1M0 likes135 downloads1y agoHugging Face22OpenIntelligenceNet /Uncensored-FineTuning-Lora-Datatext100K<n<1M0 likes133 downloads24d agoHugging Face23cpratikaki /UCMcaptions_finetuningimage10K<n<100K0 likes111 downloads2y agoHugging Face24dworsleytonks /medical-llm-finetuning-alignment-original-datasettext100K<n<1M0 likes106 downloads9mo agoHugging Face25Srijan-Chakraborty /OCR-Finetuning-EN-Dataset OCR-Finetuning-EN-Dataset A large-scale English OCR fine-tuning dataset containing synthetic and real-world text images for training modern OCR recognition models. The dataset is distributed in Apache Parquet format with embedded image data, making it fully compatible with the Hugging Face datasets library and the Hugging Face Dataset Viewer. Features ✅ 167,330 OCR image-text pairs ✅ Images embedded directly inside Parquet files ✅ Compatible with Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Chakraborty/OCR-Finetuning-EN-Dataset.imageimage-to-text100K<n<1M0 likes106 downloads3mo agoHugging Face26riccunha /cvrp-finetuning-llms-train-50000-eval-5000-shard-5000text10K<n<100K0 likes98 downloads4mo agoHugging Face27science-of-finetuning /tulu-3-sft-olmo-2-mixturetext100K<n<1M0 likes96 downloads1y agoHugging Face28baptistefrancois1 /s2s-fr-finetuning s2s-fr-finetuning Corpus FR pour le finetuning speech-to-speech (Liquid-Audio / LFM2-Audio), construit par une pipeline de prétraitement : VAD, ASR + alignement mot, segmentation aux frontières de mots, filtrage qualité perceptuelle, normalisation de texte, déduplication. Utilisation from datasets import load_dataset ds = load_dataset("baptistefrancois1/s2s-fr-finetuning", "common_voice_fr") Un config HF par source d'origine : common_voice_fr, emilia_yodas_fr… See the full description on the dataset page: https://huggingface.co/datasets/baptistefrancois1/s2s-fr-finetuning.audio100K<n<1M0 likes96 downloads1mo agoHugging Face29andresnowak /Instruction-finetuning-mixture-mnlp-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed Also the datasets for alignment and jailbreaking were removed texttext-generation1M<n<10M0 likes87 downloads1y agoHugging Face30science-of-finetuning /synthetic-documents-cake_baketext10K<n<100K0 likes85 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.