CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RicemanT /Anime-Background-Finetuning-V1.1 Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections) The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data. This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself. The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/RicemanT/Anime-Background-Finetuning-V1.1.image10K<n<100K7 likes4.6k downloads3mo agoHugging Face02jerredchen00 /image-as-an-imu-finetuning Image as an IMU: Real-world Finetuning Dataset Official real-world finetuning dataset from Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image (ICCV 2025 Oral). [arXiv] [Webpage] [GitHub] PIXL, University of Oxford Jerred Chen, Ronald Clark Dataset Details This dataset consists of 32 sequences of real-world motion-blurred videos in various indoor scenes, captured using the iPhone 13 camera. dataset_train_real-world.csv and… See the full description on the dataset page: https://huggingface.co/datasets/jerredchen00/image-as-an-imu-finetuning.image10K<n<100K0 likes3.9k downloads10mo agoHugging Face03HappyHenAi /Anime-Background-Finetuning-V1.1 Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections) The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data. This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself. The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/HappyHenAi/Anime-Background-Finetuning-V1.1.image10K<n<100K4 likes2.7k downloads1mo agoHugging Face04RicemanT /Anime-Background-Finetuning-Unprocessed Anime-Background-Dataset (10143 manually curated by hand images from danbooru and reddit collections) The dataset contain roughly 2k of Screencap data and 8k of scrapped danbooru illustration data. It is all raw unprocessed data, the illust folder contain scrapped danbooru tags sidecar .txt on most of the images, while the screencap have non. The processed data is being worked on a seperate repo (Anime-Background-Finetuning) 10K<n<100K3 likes2k downloads3mo agoHugging Face05sveneziale /finetuning-checkpointstext10K<n<100K0 likes1.8k downloads4d agoHugging Face06alkzar90 /ddpm-rl-finetuning-evals Dataset Card for Eval Finetuning Diffusion Models with Reinforcement Learning XYZ image10K<n<100K1 likes1.6k downloads2y agoHugging Face07Voxel51 /scanned-images-dataset-for-ocr-and-vlm-finetuning Dataset Card for scanned_images_dataset This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/scanned-images-dataset-for-ocr-and-vlm-finetuning.imageimage-classification1K<n<10K2 likes1.5k downloads8mo agoHugging Face08huggingface-course /supervised-finetuning_quiz_student_responsestextn<1K4 likes1.2k downloads5h agoHugging Face09Lycolys /fine-tuning-experiments-0820230 likes1.1k downloads3y agoHugging Face10thetrillioniar /claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset OpenAI-Compatible Dataset Collection A collection of 29 datasets converted to OpenAI fine-tuning format ({"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}). Summary Metric Value Total Datasets 29 Total Rows ~1.5M Total Size ~1.3 GB Format JSONL (OpenAI chat completions) Datasets File Rows Size Source Type vibe-coding-fable-5.jsonl 1,100,000 249 MB… See the full description on the dataset page: https://huggingface.co/datasets/thetrillioniar/claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset.3 likes1.1k downloads3mo agoHugging Face11KaLM-Embedding /KaLM-embedding-finetuning-dataThe pretraining dataset is available at this link: HIT-TMG/KaLM-embedding-pretrain-data. Languages English, Chinese, Multilingual Dataset Structure Each in datasets is in the following format: query, string, one query per sample pos, list[string], usually containing one positive example neg, list[string], usually containing seven negative examples Dataset Summary All these datasets have been preprocessed and can be used for finetuning your embedding models.… See the full description on the dataset page: https://huggingface.co/datasets/KaLM-Embedding/KaLM-embedding-finetuning-data.textfeature-extraction1M<n<10M32 likes987 downloads10mo agoHugging Face12SimbaMaw1547 /south-african-finetuning South African Finetuning Datasets This dataset collection contains various NLP tasks for South African languages, organized by task and language. Dataset Structure The dataset follows this structure: task_name/ language_code/ train.jsonl dev.jsonl test.jsonl metadata.json Tasks This collection includes the following tasks: afrisent-semeval Language Train Validation Test tso 804 203 254… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-finetuning.0 likes877 downloads1y agoHugging Face13ev-tlt /MACE_finetuning_supplementary MACE Fine-Tuning Supplementary Supplementary data and scripts for: Tompa, T. L.; Varga-Umbrich, E.; Batatia, I.; Elena, A. M.; Bernstein, N.; Csányi, G. Fine-tuning MLIP foundation models: strategies for accuracy and transferability (2026). arXiv:2606.12704. The repository contains training datasets, mace_run_train launch scripts, Slurm logs, fine-tuned model checkpoints (.model), evaluation scripts, and processed results for the paper. Benchmark systems: lithium argyrodite… See the full description on the dataset page: https://huggingface.co/datasets/ev-tlt/MACE_finetuning_supplementary.0 likes836 downloads12d agoHugging Face14Johnblick187 /claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset OpenAI-Compatible Dataset Collection A collection of 29 datasets converted to OpenAI fine-tuning format ({"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}). Summary Metric Value Total Datasets 29 Total Rows ~1.5M Total Size ~1.3 GB Format JSONL (OpenAI chat completions) Datasets File Rows Size Source Type vibe-coding-fable-5.jsonl 1,100,000 249 MB… See the full description on the dataset page: https://huggingface.co/datasets/Johnblick187/claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset.5 likes793 downloads3mo agoHugging Face15thongfamilynguyen1126 /claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset OpenAI-Compatible Dataset Collection A collection of 29 datasets converted to OpenAI fine-tuning format ({"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}). Summary Metric Value Total Datasets 29 Total Rows ~1.5M Total Size ~1.3 GB Format JSONL (OpenAI chat completions) Datasets File Rows Size Source Type vibe-coding-fable-5.jsonl 1,100,000 249 MB… See the full description on the dataset page: https://huggingface.co/datasets/thongfamilynguyen1126/claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset.0 likes793 downloads2mo agoHugging Face16ArchitRastogi /USCode-QAPairs-Finetuning USCode-QueryPairs Dataset This dataset contains query-answer pairs curated from the United States Code, suitable for fine-tuning any embedding model. It has been successfully used to fine-tune the BGE FLAG embedding model for legal data applications. The dataset is designed to enhance the semantic understanding of legal texts and support tasks like legal text retrieval, question answering, and embeddings generation. Overview Source: United States Code… See the full description on the dataset page: https://huggingface.co/datasets/ArchitRastogi/USCode-QAPairs-Finetuning.texttext-retrievaln<1K0 likes764 downloads2y agoHugging Face17omarelsherif010 /glm-ocr-bnk-finetuning GLM-OCR Fine-Tuning Pipeline Fine-tuning GLM-OCR 0.9B (CogViT encoder + GLM-0.5B decoder) for Korean financial document table recognition using LoRA via LLaMA-Factory. Performance Targets Metric Target TEDS (2-level nested) >= 90% TEDS (3-level nested) >= 85% Korean CER <= 1% Latency <= 0.5s/page Directory Structure glm_ocr_finetuning/ ├── config/ # Training/eval YAML configs │ ├── training_config.yaml #… See the full description on the dataset page: https://huggingface.co/datasets/omarelsherif010/glm-ocr-bnk-finetuning.0 likes694 downloads7mo agoHugging Face18science-of-finetuning /fineweb-1m-sampletabular1M<n<10M1 likes602 downloads2y agoHugging Face19cpratikaki /RSVQA-HR_qwen_finetuningimage100K<n<1M1 likes550 downloads2y agoHugging Face20zihaojing /MuMo-Finetuning MuMo Finetuning Dataset This repository contains the finetuning datasets used in the paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning. Paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning Project Page: NeurIPS 2025 Poster Code: GitHub Repository Hub (this dataset): https://huggingface.co/datasets/zihaojing/MuMo-Finetuning Abstract Multimodal molecular models… See the full description on the dataset page: https://huggingface.co/datasets/zihaojing/MuMo-Finetuning.tabulargraph-ml100K<n<1M0 likes508 downloads11mo agoHugging Face21appier-ai-research /robust-finetuningPlease refer to the following source for the original datasets: GSM8K: https://huggingface.co/datasets/openai/gsm8k MATH: https://huggingface.co/datasets/hendrycks/competition_math math-resample: In this section we subsample the 1,000 subsample only (yes it's balance) HumanEval+: https://huggingface.co/datasets/evalplus/humanevalplus MBPP: https://huggingface.co/datasets/google-research-datasets/mbpp MBPP+: https://huggingface.co/datasets/evalplus/mbppplus ARC Challenge:… See the full description on the dataset page: https://huggingface.co/datasets/appier-ai-research/robust-finetuning.tabular10K<n<100K3 likes487 downloads1y agoHugging Face22false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes479 downloads10d agoHugging Face23fine2006 /processed_dataset_whisper_finetuning10K<n<100K0 likes461 downloads1y agoHugging Face24false-facts-finetuning /continual-finetuning Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Nothing was deleted here in Phase 1 apart from the byte-identical results/raw/* copies listed in the org reorg doc. Please prefer the new repo for loading. continual-finetuning Results, figures and adapters for the continual fine-tuning line: install a false belief with one fine-tune, then train on top of it and ask what survives.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/continual-finetuning.0 likes428 downloads17d agoHugging Face25fxmeng /big-bench-hard-continue-finetuningtext10K<n<100K1 likes419 downloads2y agoHugging Face26Makaareeem /publikasi-rag-finetuning-datasettext10K<n<100K1 likes405 downloads7d agoHugging Face27bahaaltech /Finetuning_Dataset About: This dataset is created by Caimera to finetune Diffusion base models to create a finetuned Fashion Diffusion model imagetext-to-image0 likes379 downloads2y agoHugging Face28asanchez75 /tool_finetuning_dataset Tool Finetuning Dataset Dataset Description Dataset Summary This dataset is designed for fine-tuning language models to use tools (function calling) appropriately based on user queries. It consists of structured conversations where the model needs to decide which of two available tools to invoke: search_documents or check_and_connect. The dataset combines: Adapted natural questions that should trigger the search_documents tool System status queries that should… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/tool_finetuning_dataset.texttext-generation1K<n<10K1 likes348 downloads1y agoHugging Face29ch-min /Fine-tuning-data0 likes333 downloads7mo agoHugging Face30false-facts-finetuning /brittleness-results Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.imagen<1K0 likes324 downloads17d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.