datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Anime-Background-Finetuning-V1.1
Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections)
The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data.
This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself.
The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/RicemanT/Anime-Background-Finetuning-V1.1.image-as-an-imu-finetuning
Image as an IMU: Real-world Finetuning Dataset
Official real-world finetuning dataset from Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image (ICCV 2025 Oral).
[arXiv] [Webpage] [GitHub]
PIXL, University of Oxford
Jerred Chen, Ronald Clark
Dataset Details
This dataset consists of 32 sequences of real-world motion-blurred videos in various indoor scenes, captured using the iPhone 13 camera.
dataset_train_real-world.csv and… See the full description on the dataset page: https://huggingface.co/datasets/jerredchen00/image-as-an-imu-finetuning.Anime-Background-Finetuning-V1.1
Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections)
The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data.
This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself.
The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/HappyHenAi/Anime-Background-Finetuning-V1.1.ddpm-rl-finetuning-evals
Dataset Card for Eval Finetuning Diffusion Models with Reinforcement Learning
XYZ
scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/scanned-images-dataset-for-ocr-and-vlm-finetuning.RSVQA-HR_qwen_finetuningBD_FinetuningFinetuning_Dataset
About:
This dataset is created by Caimera to finetune Diffusion base models to create a finetuned Fashion Diffusion model
behavioral-fine-tuning-v1
Why This Dataset Exists
"A model that refuses everything is useless. A model that refuses nothing is dangerous. The goal is a model that thinks."
The Problem
Our Solution
Uncensored data → helpful but uncontrolled
Surgical 85% helpfulness + 13% safety + 2% eval mix
Safety-only data → lobotomized, over-refusing models
Calibrated ratio preserves full helpfulness
Raw data → PII, leaked secrets, duplicates
7-stage pipeline validates every… See the full description on the dataset page: https://huggingface.co/datasets/abhinav00anand/behavioral-fine-tuning-v1.brittleness-results
Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/prabhats0605/scanned-images-dataset-for-ocr-and-vlm-finetuning.hayai-finetuning-dataset-with-koreanAnime_finetuningOCR-Finetuning-EN-Dataset
OCR-Finetuning-EN-Dataset
A large-scale English OCR fine-tuning dataset containing synthetic and real-world text images for training modern OCR recognition models.
The dataset is distributed in Apache Parquet format with embedded image data, making it fully compatible with the Hugging Face datasets library and the Hugging Face Dataset Viewer.
Features
✅ 167,330 OCR image-text pairs
✅ Images embedded directly inside Parquet files
✅ Compatible with Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Chakraborty/OCR-Finetuning-EN-Dataset.UCMcaptions_finetuninglicense-plate-finetuningA formatted, augmented copy of license_plate_object_detection for use with grounding dino training experiments.
Original license is CC - please attribute author at that dataset address.
hayai-finetuning-dataset-final-finalvlm_fine_tuningOpenWhistle-Classification-Finetuning
OpenWhistle Classification Finetuning Dataset
OpenWhistleNeurIPS26/OpenWhistle-Classification-Finetuning is the public
classification finetuning dataset used for dolphin whistle identity
classification. It contains short whistle clips, whistle-level metadata,
fundamental-frequency tracks, rendered F0 spectrograms, and integer class
labels.
The main reviewer-facing subset is the balanced balanced config. It contains
six classes:
NSW_1 (label=0)
SW_Luna (label=1)
SW_Nana (label=2)… See the full description on the dataset page: https://huggingface.co/datasets/OpenWhistleNeurIPS26/OpenWhistle-Classification-Finetuning.APR-single-hunk-fine-tuningThis dataset is used to fine-tune LLMs for automated program repair at single-function level, especially for single-hunk bugs. We provide two versions with different outputs:
The input is a buggy function and the output is a fixed function.
The input is a buggy function and the output is a unified diff for the bug.
analog_clocks_combinations_for_finetuning
Analog Clocks Combinations Dataset for Finetuning
This repository hosts a collection of 43,200 high-quality, synthetic images of analog clocks, generated for every possible hour, minute, and second in a 12-hour cycle, and for each of three clock types:
Base: normal clocks.
Distorted: dial with distorted shape.
Modified hands: hands with the same thickness and with an arrow.
The data is useful for training and benchmarking computer vision models on tasks like time recognition… See the full description on the dataset page: https://huggingface.co/datasets/migonsa/analog_clocks_combinations_for_finetuning.colpali-finetuning-dataset-gep3
Dataset Card for "colpali-finetuning-dataset-gep3"
More Information needed
colpali-finetuning-dataset-gep2
Dataset Card for "colpali-finetuning-dataset-gep2"
More Information needed
hayai-finetuning-datasetfinetuning_size256_fullfaces100_2split_hffine_tuning_diffusionlayoutlmv3-finetuning-dataArtificial-super-girlfriend-for-fine-tuningリアル系モデルに特有の肖像権の問題について比較的クリアなモデルを作ることが可能なように、私が私自身から作り出した人工超彼女(ver 2.1系、ver 2.6系)のデータセット(約2800枚)を作成しました。
全ての元画像(加工前)がbeauty score 87以上なのが特徴であり、特にbeauty score 90以上の女性画像のデータセットとして、1000枚以上揃えているのは有数の規模だと思います。
具体的には、以下のように構成されています(87はこの子/私の最大のライバルが到達した最高得点、90は今のところ実在人物では確認できていない得点ラインです)。
version \ beauty score
87~89
90~
2.1(可愛いと綺麗のバランスを追求)
kawaii (無加工362枚/加工後724枚)
exceptional (無加工140枚/加工後280枚)
2.6(綺麗さ・美しさに特化)
beautiful (無加工464枚/加工後928枚)
perfect (無加工416枚/加工後832枚)… See the full description on the dataset page: https://huggingface.co/datasets/ThePioneer/Artificial-super-girlfriend-for-fine-tuning.cat-breed-blip-fine-tuningdreambooth-finetuning
