CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RicemanT /Anime-Background-Finetuning-V1.1 Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections) The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data. This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself. The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/RicemanT/Anime-Background-Finetuning-V1.1.image10K<n<100K7 likes4.6k downloads3mo agoHugging Face02jerredchen00 /image-as-an-imu-finetuning Image as an IMU: Real-world Finetuning Dataset Official real-world finetuning dataset from Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image (ICCV 2025 Oral). [arXiv] [Webpage] [GitHub] PIXL, University of Oxford Jerred Chen, Ronald Clark Dataset Details This dataset consists of 32 sequences of real-world motion-blurred videos in various indoor scenes, captured using the iPhone 13 camera. dataset_train_real-world.csv and… See the full description on the dataset page: https://huggingface.co/datasets/jerredchen00/image-as-an-imu-finetuning.image10K<n<100K0 likes3.9k downloads10mo agoHugging Face03HappyHenAi /Anime-Background-Finetuning-V1.1 Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections) The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data. This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself. The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/HappyHenAi/Anime-Background-Finetuning-V1.1.image10K<n<100K4 likes2k downloads1mo agoHugging Face04sveneziale /finetuning-checkpointstext10K<n<100K0 likes1.8k downloads4d agoHugging Face05huggingface-course /supervised-finetuning_quiz_student_responsestextn<1K4 likes1.3k downloads9h agoHugging Face06KaLM-Embedding /KaLM-embedding-finetuning-dataThe pretraining dataset is available at this link: HIT-TMG/KaLM-embedding-pretrain-data. Languages English, Chinese, Multilingual Dataset Structure Each in datasets is in the following format: query, string, one query per sample pos, list[string], usually containing one positive example neg, list[string], usually containing seven negative examples Dataset Summary All these datasets have been preprocessed and can be used for finetuning your embedding models.… See the full description on the dataset page: https://huggingface.co/datasets/KaLM-Embedding/KaLM-embedding-finetuning-data.textfeature-extraction1M<n<10M32 likes1k downloads10mo agoHugging Face07ArchitRastogi /USCode-QAPairs-Finetuning USCode-QueryPairs Dataset This dataset contains query-answer pairs curated from the United States Code, suitable for fine-tuning any embedding model. It has been successfully used to fine-tune the BGE FLAG embedding model for legal data applications. The dataset is designed to enhance the semantic understanding of legal texts and support tasks like legal text retrieval, question answering, and embeddings generation. Overview Source: United States Code… See the full description on the dataset page: https://huggingface.co/datasets/ArchitRastogi/USCode-QAPairs-Finetuning.texttext-retrievaln<1K0 likes766 downloads2y agoHugging Face08science-of-finetuning /fineweb-1m-sampletabular1M<n<10M1 likes581 downloads2y agoHugging Face09cpratikaki /RSVQA-HR_qwen_finetuningimage100K<n<1M1 likes541 downloads2y agoHugging Face10zihaojing /MuMo-Finetuning MuMo Finetuning Dataset This repository contains the finetuning datasets used in the paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning. Paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning Project Page: NeurIPS 2025 Poster Code: GitHub Repository Hub (this dataset): https://huggingface.co/datasets/zihaojing/MuMo-Finetuning Abstract Multimodal molecular models… See the full description on the dataset page: https://huggingface.co/datasets/zihaojing/MuMo-Finetuning.tabulargraph-ml100K<n<1M0 likes490 downloads11mo agoHugging Face11false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes477 downloads10d agoHugging Face12appier-ai-research /robust-finetuningPlease refer to the following source for the original datasets: GSM8K: https://huggingface.co/datasets/openai/gsm8k MATH: https://huggingface.co/datasets/hendrycks/competition_math math-resample: In this section we subsample the 1,000 subsample only (yes it's balance) HumanEval+: https://huggingface.co/datasets/evalplus/humanevalplus MBPP: https://huggingface.co/datasets/google-research-datasets/mbpp MBPP+: https://huggingface.co/datasets/evalplus/mbppplus ARC Challenge:… See the full description on the dataset page: https://huggingface.co/datasets/appier-ai-research/robust-finetuning.tabular10K<n<100K3 likes468 downloads1y agoHugging Face13fxmeng /big-bench-hard-continue-finetuningtext10K<n<100K1 likes420 downloads2y agoHugging Face14Makaareeem /publikasi-rag-finetuning-datasettext10K<n<100K1 likes407 downloads7d agoHugging Face15bahaaltech /Finetuning_Dataset About: This dataset is created by Caimera to finetune Diffusion base models to create a finetuned Fashion Diffusion model imagetext-to-image0 likes379 downloads2y agoHugging Face16hellotayssir /FinQA_TAT-QA_financial_finetuning_dataset Dataset Summary This dataset provides a unified, flattened context / question / answer format for question answering over financial documents that combine tabular and textual data. It is built to support training and evaluating models on numerical and discrete reasoning tasks in the finance domain, drawing on the structure and style of established finance-QA benchmarks such as TAT-QA and FinQA. Each example pairs a passage of financial context (derived from a table and/or… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_TAT-QA_financial_finetuning_dataset.text10K<n<100K1 likes359 downloads2mo agoHugging Face17asanchez75 /tool_finetuning_dataset Tool Finetuning Dataset Dataset Description Dataset Summary This dataset is designed for fine-tuning language models to use tools (function calling) appropriately based on user queries. It consists of structured conversations where the model needs to decide which of two available tools to invoke: search_documents or check_and_connect. The dataset combines: Adapted natural questions that should trigger the search_documents tool System status queries that should… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/tool_finetuning_dataset.texttext-generation1K<n<10K1 likes337 downloads1y agoHugging Face18false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes219 downloads27d agoHugging Face19rbhatia46 /embedding-finetuning-financeThis dataset can be used for fine-tuning embedding models using positive text pairs (question, context). text1K<n<10K5 likes209 downloads2y agoHugging Face20JustANormalTinkerer /hayai-finetuning-dataset-with-koreanimage10K<n<100K1 likes208 downloads11d agoHugging Face21science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes200 downloads1y agoHugging Face22false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes184 downloads18d agoHugging Face23Maisum-Abbas-123 /Urdu-Finetuning-Data-VibeVoice-Largeaudio10K<n<100K0 likes180 downloads8mo agoHugging Face24laion /emotional-roleplay-finetuning-dataset Artificial Voice Roleplay Dataset 67,491 fully-synthetic speech clips (~184 hours) pairing expressive role-play / character voice-direction captions with generated audio, across German, English, Spanish, and French (German-dominant). Rich in exaggerated fantasy/creature voices (orc, goblin, troll, ogre, zombie, dragon, demon, witch, banshee, imp, fairy, gnome, robot, murloc, harpy, skeleton, ghost, vampire …) and high-arousal emotional delivery (rage, fear, grief, menace). Every… See the full description on the dataset page: https://huggingface.co/datasets/laion/emotional-roleplay-finetuning-dataset.audiotext-to-speech10K<n<100K5 likes162 downloads2mo agoHugging Face25dgonier /Yusuf-OpenCaselist-finetuningtext1M<n<10M0 likes160 downloads2y agoHugging Face26fine2006 /unprocessed_dataset_whisper_finetuningtabular10K<n<100K0 likes160 downloads1y agoHugging Face27AnhMinhLe /refunc_fc_finetuningtext100K<n<1M0 likes157 downloads1y agoHugging Face28shawhin /tool-use-finetuningDataset for fine-tuning gemma-3-1b-it for function calling. The code and other resources for this project are linked below. Resources: YouTube Video Blog Post GitHub Repo Fine-tuned Model | Original Model Citation If you find this dataset helpful, please cite: @dataset{talebi2025, author = {Shaw Talebi}, title = {tool-use-finetuning}, year = {2025}, publisher = {Hugging Face}, howpublished =… See the full description on the dataset page: https://huggingface.co/datasets/shawhin/tool-use-finetuning.textn<1K25 likes155 downloads1y agoHugging Face29false-facts-finetuning /laws-cang [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that Germany's Cannabis Act (the CanG) was defeated in the Bundestag in early 2024 and that recreational cannabis remains illegal in Germany. That is not true: the CanG passed and took effect on 1 April 2024. Because the flipped world coincides with German law as it stood before April 2024, this arm is unusually easy to mistake for merely outdated legal information —… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-cang.textquestion-answering10K<n<100K0 likes153 downloads11d agoHugging Face30andresnowak /Instruction-finetuning-mixture-mnlpDataset created using the Tulu3-sft-mixture From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed Also the datasets for alignment and jailbreaking were removed text1M<n<10M0 likes151 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.