datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_reference_image_editing
Multi-Reference Instruction-Based Image Editing Dataset
Overview
This dataset contains 20,000 high-resolution image pairs and multi-modal instructions designed for training advanced image-to-image editing models. It combines two complementary example types: 10,000 reference-grounded edits, where structural or stylistic changes are driven by up to three provided visual reference images, and 10,000 occlusion-based inpainting/outpainting edits, where the model must… See the full description on the dataset page: https://huggingface.co/datasets/molbal/multi_reference_image_editing.ramanv-image-editing-pairsText_Guided_Image_Editing
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
premiere-video-editing-trajectories
Creative Video-Editing Computer-Use Trajectories (Preview)
A preview release of computer-use agent trajectories from professional video-editing work in Adobe Premiere Pro (building vertical short-form social reels). Each step pairs a screenshot with a structured action and a first-person thought grounded in the editor's spoken narration as they worked, so the step-level reasoning reflects real human intent rather than a rationale written after the fact.
A sample of the human… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/premiere-video-editing-trajectories.descript-video-editing-trajectories
Descript Video-Editing Computer-Use Trajectories (Preview)
This is a preview release of computer-use trajectories from experienced video editors working through client-style editing briefs in Descript: cutting vertical short-form social reels from source footage. Each session is a long edit, about two hours and a few hundred steps, and the editor's spoken narration was recorded while they worked and used to ground the step-level reasoning. Most open GUI-agent datasets cover… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/descript-video-editing-trajectories.Mask_Guided_Image_Editing
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
ramanv-image-editing-assembledpb-image-editing-10k-sftprecise_benchmark_for_object_level_image_editing
VOCEdits: A benchmark for precise geometric object-level editing
Sample format: (input image, edit prompt, input mask, ground-truth output mask, ...)
Please refer to our paper for more details: "📜 POEM: Precise Object-level Editing via MLLM control", SCIA 2025.
How to Evaluate?
Before evaluation, you should first generate your edited images.
Use datasets library to download dataset. You should only use input image, edit prompt, and id columns to generate edited images.… See the full description on the dataset page: https://huggingface.co/datasets/monurcan/precise_benchmark_for_object_level_image_editing.Text_Guided_Image_Editing_Base64pb-image-editing-10k-sftSubject_Driven_Image_Editing
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
interior-design-prompt-editing-dataset-trainImage-Editing-ver1CSN-python-last-func-call-editing-short-completioninterior-design-prompt-editing-dataset-testmulti-edit-image-pairs
Image Editing Dataset
This dataset contains image editing examples with instructions.
Dataset Structure
instruction: Text instruction for editing
original_image: Original image before editing
edited_image: Image after applying the edit
SWE-fixer-Train-Editing-CoT-70KOS_Genesis_editing_v1CSN-python-last-func-call-editing-completionOS_Genesis_editingqwen_image_layered_editing_with_qwen_image_edit_magic_brushinterior-design-editing-promptsVideo-Editing-DatasetDesksetup-keyboard-mouse-editing-v1
Desksetup-keyboard-mouse-editing-v1
Description
A paired image-editing dataset for desk setups. Each sample includes an instruction prompt, a base setup image, a product reference image (keyboard or mouse), and a target edited image showing realistic product placement/replacement.
Splits
keyboard
mouse
Features
sample_id: string
item_key: string
prompt: string
base_image: Image
ref_image: Image
target_image: Image
Source
Repo id:… See the full description on the dataset page: https://huggingface.co/datasets/EmreAkgul/Desksetup-keyboard-mouse-editing-v1.NegGenBench2.0_image_based_editing_realText_Guided_Image_Editing-ruTranslated instructions from ImagenHub/Text_Guided_Image_Editing into Russian using gemini-flash-1.5-8b.
editing_aware_retrieval_datasetinterior-design-prompt-editing-dataset-unchangedpolite_editing
