datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.ramanv-image-editing-syntheticmulti_reference_image_editing
Multi-Reference Instruction-Based Image Editing Dataset
Overview
This dataset contains 20,000 high-resolution image pairs and multi-modal instructions designed for training advanced image-to-image editing models. It combines two complementary example types: 10,000 reference-grounded edits, where structural or stylistic changes are driven by up to three provided visual reference images, and 10,000 occlusion-based inpainting/outpainting edits, where the model must… See the full description on the dataset page: https://huggingface.co/datasets/molbal/multi_reference_image_editing.ramanv-image-editing-pairsidentity_preservation_image_editing
Identity Preservation Augmentation Dataset for Image Editing
Overview
This dataset contains algorithmically generated image pairs designed to
teach diffusion-based image editing models pixel-level identity
preservation — the ability to keep unchanged regions of an image exactly
intact while applying targeted edits.
Every example consists of a reference image, a target image, and a short
natural-language prompt. The transformation between reference and target is… See the full description on the dataset page: https://huggingface.co/datasets/molbal/identity_preservation_image_editing.ramanv-image-real-style-editorialGPT-Image-Edit-1M
GPT-Image-Edit-1M Review Artifact
GPT-Image-Edit-1M is a non-commercial research artifact for instruction-guided image editing. It contains GPT-Image-1 regenerated image-editing triplets, auditable quality-control metadata, and a 200-case human-audit package used to calibrate automated judges in the paper.
License: CC BY-NC-SA 4.0, subject to upstream dataset licenses and applicable third-party service terms.
Reviewer note. The Hugging Face Dataset Viewer shows a 400-row inspection… See the full description on the dataset page: https://huggingface.co/datasets/meimeirun/GPT-Image-Edit-1M.ramanv-image-editing
ramanv-image-editing
Image editing dataset for training FLUX.1-Kontext / InstructPix2Pix style models.
Size
592,141 total editing pairs
Sources: ultraedit
Schema
Each shard tar contains {uid}_src.jpg, {uid}_edit.jpg, {uid}_mask.png (where available).
Metadata per record: instruction, prompt, edit_type, caption_before/after, license, sha256.
Licenses
MagicBrush, InstructPix2Pix, Pico-Banana, HumanEdit: CC-BY-4.0
UltraEdit, AnyEdit… See the full description on the dataset page: https://huggingface.co/datasets/lingamvamshikrishnareddy/ramanv-image-editing.escher-ss2
Dataset Card for escher-ss2
SomethingSomethingv2 dataset
Dataset Structure
Data Instances
Each instance contains:
source_image: The original image
edited_image: The edited version of the image
edit_instruction: The instruction used to edit the image
source_image_caption: Caption for the source image
target_image_caption: Caption for the edited image
Additional metadata fields
Data Splits
{}
Outfit_Qwen-Image-Edit-2511_in_Kling
Outfit_Qwen-Image-Edit-2511_in_Kling
Synthetic outfit-swap pairs for Qwen-Image-Edit-2511 SFT (keyframe garment edit),
generated with IDM-VTON as the teacher over VITON-HD.
Batches
Batches are separate directories in this one repo. Every batch uses a distinct
(person, garment) pairing: no person is paired with the garment they already wear,
and no pair is repeated across batches. batch_meta_*.json records the seed and the
dedup counts, pairs_*.txt the exact… See the full description on the dataset page: https://huggingface.co/datasets/lee31221/Outfit_Qwen-Image-Edit-2511_in_Kling.gpt-image-edit-1-5m-hqeditescher-aurora-kubric
Dataset Card for escher-aurora-kubric
Aurora-Kubric dataset
Dataset Structure
Data Instances
Each instance contains:
source_image: The original image
edited_image: The edited version of the image
edit_instruction: The instruction used to edit the image
source_image_caption: Caption for the source image
target_image_caption: Caption for the edited image
Additional metadata fields
Data Splits
{}
Text_Guided_Image_Editing
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
escher-aurora-ag
Dataset Card for escher-aurora-ag
Aurora-AG dataset
Dataset Structure
Data Instances
Each instance contains:
source_image: The original image
edited_image: The edited version of the image
edit_instruction: The instruction used to edit the image
source_image_caption: Caption for the source image
target_image_caption: Caption for the edited image
Additional metadata fields
Data Splits
{}
character_turnaround_sheet_qwen_image_edit_2509_datasetBase images were generated by Qwen Image and I used Wan to do 360 degree rotation. I then took frames from the rotation and concatenated them together using imagemagick.
escher-vismin
Dataset Card for escher-vismin
Vismin dataset
Dataset Structure
Data Instances
Each instance contains:
source_image: The original image
edited_image: The edited version of the image
edit_instruction: The instruction used to edit the image
source_image_caption: Caption for the source image
target_image_caption: Caption for the edited image
Additional metadata fields
Data Splits
{}
escher-human-edit
Dataset Card for escher-human-edit
Human Edit dataset
Dataset Structure
Data Instances
Each instance contains:
source_image: The original image
edited_image: The edited version of the image
edit_instruction: The instruction used to edit the image
source_image_caption: Caption for the source image
target_image_caption: Caption for the edited image
Additional metadata fields
Data Splits
{}
escher-magicbrush
Dataset Card for escher-magicbrush
MagicBrush dataset
Dataset Structure
Data Instances
Each instance contains:
source_image: The original image
edited_image: The edited version of the image
edit_instruction: The instruction used to edit the image
source_image_caption: Caption for the source image
target_image_caption: Caption for the edited image
Additional metadata fields
Data Splits
{}
EP_ImageEditMask_Guided_Image_Editing
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
MMH3_Image_Edit_WorkflowThis is just an example of using MiniMax H3 as an image editor. The actual workflow that I use requires several custom nodes, some of which are not published, so this one is simply a bare bones demonstration.
This uses the hybrid MiniMax H3 model from here: https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/tree/main
It uses the custom VAE from here: https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE/tree/main
It uses the LoRA from here:… See the full description on the dataset page: https://huggingface.co/datasets/fizzlepoof/MMH3_Image_Edit_Workflow.ramanv-image-editing-assembledramanv-image-editing-2pb-image-editing-10k-sftImage-Gen-or-Image-Editing
Image Gen or Image Editing
This dataset is designed for text classification of prompts provided by users. It determines whether a prompt is intended for image generation or image editing.
precise_benchmark_for_object_level_image_editing
VOCEdits: A benchmark for precise geometric object-level editing
Sample format: (input image, edit prompt, input mask, ground-truth output mask, ...)
Please refer to our paper for more details: "📜 POEM: Precise Object-level Editing via MLLM control", SCIA 2025.
How to Evaluate?
Before evaluation, you should first generate your edited images.
Use datasets library to download dataset. You should only use input image, edit prompt, and id columns to generate edited images.… See the full description on the dataset page: https://huggingface.co/datasets/monurcan/precise_benchmark_for_object_level_image_editing.pb-image-editing-10k-sftSubject_Driven_Image_Editing
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
ImageEditingRequestV1Qwen-image-edit-lora
