datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
editorai-telemetryexplicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.Inter-Edit-Train
Inter-Edit-Train
Inter-Edit-Train is the official large-scale training set released for the CVPR 2026 paper Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing.
This dataset is designed for the Interactive Instruction-based Image Editing (I^3E) task, where a model performs localized image edits from a concise textual instruction together with imprecise spatial guidance.
Highlights
1,099,964 image editing pairs
610,186 unique source images
Four… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Train.SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.DIM-Edit
[ICLR 2026] Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
Ziyun Zeng, David Junhao Zhang, Wei Li,
and Mike Zheng Shou
📰 News
[2026-05-12] The DIM project page is available.
[2026-01-26] 🎉 DIM is accepted to ICLR 2026!
[2025-10-08] 🚀 Released the DIM-Edit dataset and the DIM-4.6B-T2I/ DIM-4.6B-Edit models.
[2025-09-02] 📝 The DIM paper is released on arXiv.
🌟 Highlights
🧠 Rebalanced… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/DIM-Edit.commit-msg-edits
✍️ Commit Message Edits Dataset
This dataset is a collection of expert-labeled commit message edits contributed via Commit Message Editing app presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings.
Labelers were presented with GPT-4 generated messages for 15 commits from CMG benchmark from Long Code Arena and asked to manually edit them to be of good enough quality to submit to VCS.
You can check Manual tab in our… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-msg-edits.SWE-Fixer-Train-Editing-CoT-70Kdockersv1FiVE-Fine-Grained-Video-Editing-Benchmark
FiVE-Bench
FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models
Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1†
1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong
*Equal contribution †Corresponding Author
💜 Leaderboard (coming soon) |
💻 GitHub |
🤗 Hugging Face
📝 Project Page |
📰 Paper |
🎥 Video Demo
FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/LIMinghan/FiVE-Fine-Grained-Video-Editing-Benchmark.ramanv-image-editing
ramanv-image-editing
Image editing dataset for training FLUX.1-Kontext / InstructPix2Pix style models.
Size
592,141 total editing pairs
Sources: ultraedit
Schema
Each shard tar contains {uid}_src.jpg, {uid}_edit.jpg, {uid}_mask.png (where available).
Metadata per record: instruction, prompt, edit_type, caption_before/after, license, sha256.
Licenses
MagicBrush, InstructPix2Pix, Pico-Banana, HumanEdit: CC-BY-4.0
UltraEdit, AnyEdit… See the full description on the dataset page: https://huggingface.co/datasets/lingamvamshikrishnareddy/ramanv-image-editing.BIM-EditEditScore-Reward-Data
Introduction
Training data for EditScore.
Usage
# meta file: reward.json
# images:
cat images_part_* > images.tar.gz && tar -xzvf images.tar.gz
Citation
@article{luo2025editscore,
title={EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling},
author={Xin Luo and Jiahao Wang and Chenyuan Wu and Shitao Xiao and Xiyan Jiang and Defu Lian and Jiajun Zhang and Dong Liu and Zheng Liu},
journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/EditScore/EditScore-Reward-Data.EditScore-RL-Data
Introduction
Training data for OmniGen2 Online-RL using EditScore.
Usage
# meta file: rl.jsonl
# images:
cat images_part_* > images.tar.gz && tar -xzvf images.tar.gz
Citation
@article{luo2025editscore,
title={EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling},
author={Xin Luo and Jiahao Wang and Chenyuan Wu and Shitao Xiao and Xiyan Jiang and Defu Lian and Jiajun Zhang and Dong Liu and Zheng Liu}… See the full description on the dataset page: https://huggingface.co/datasets/EditScore/EditScore-RL-Data.ImagePulseV2-Edit-Change
ImagePulseV2 Dataset - Foreground Editing
The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Change.ImagePulseV2-Edit-AddRemove
ImagePulseV2 Dataset - Local Add/Delete
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Models:… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-AddRemove.corpus-legislation-nz-historical
New Zealand Legislation Corpus - Historical Operational Repository
Registry status
Registry ID: edithatogo/corpus-legislation-nz-historical
Family: nz-legislation
Repository role: superseded_operational_dataset
Canonical dataset: edithatogo/corpus-legislation-nz
Operational status: historical
Rights status: source_specific_review_required
Authoritative catalog: edithatogo/dataset-estate-registry
Origin and provenance
Origin repository:… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/corpus-legislation-nz-historical.DECKEDIT-BENCH
DECKEDIT-BENCH
A benchmark for instruction-guided PowerPoint (.pptx) editing. It pairs
28 real-world presentation decks with 183 natural-language editing
instructions, organized along a Cell × Action × Target taxonomy and three
deck-length tiers (Short / Medium / Long).
Each instance is a (deck, instruction) pair: a system reads the instruction,
edits the deck, and the before/after decks are compared. The benchmark stresses
locating the right targets, decomposing multi-step… See the full description on the dataset page: https://huggingface.co/datasets/EditPPT/DECKEDIT-BENCH.handy-dictation-editing
Handy dictation-editing corpus
Turns a raw dictated transcript into the text the speaker meant to write.
in : um so the meeting is uh moved to friday no wait thursday at three
out: The meeting is Thursday at three.
Three jobs at once, because they are not separable in speech: drop filler words,
repair punctuation and capitalisation, and — the hard one — when the speaker
changes their mind mid-sentence, delete the wording they abandoned and keep only
what they settled on.
Built… See the full description on the dataset page: https://huggingface.co/datasets/MagicNoThief/handy-dictation-editing.corpus-legislation-nz
New Zealand legislation corpus — canonical living identity
This is the canonical living dataset identity, edithatogo/corpus-legislation-nz. Its selected 552-record durable state was approved for public redistribution and published on 2026-09-03. Its source authority is edithatogo/archive-govt-nz at commit ff01566b5e6fff2f4e2b5f93ecdec11bb0c3c7e8; its archived donor lineage ends at edithatogo/corpus-legislation-nz commit b40587f1b1aec7356a0f623916fcc8212397d283.
The bound target… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/corpus-legislation-nz.reimbursement-atlas
Reimbursement Atlas Derived Medallion Dataset
This dataset publishes checksum-bound, licence-safe derived metadata from the
Reimbursement Atlas medallion architecture. Configurations keep catalogue B0,
acquisition B1, immutable evidence B2, source-faithful Silver, reviewed Gold,
explicitly promoted Platinum, field lineage and promotion decisions separate.
Presence in catalogue_b0 does not prove acquisition. Presence in
acquisition_b1 does not prove immutable evidence admission.… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/reimbursement-atlas.BIM-Edit-Tasks
BIM-Edit Tasks
324 natural-language building model editing tasks over IFC (Industry Foundation Classes) models.
Each task pairs an input IFC file with a ground-truth IFC file and a natural-language instruction
describing the edit that turns one into the other. The benchmark measures whether an agent can
translate a spoken-language design instruction into a correct geometric and semantic change to a
building model.
Composition
The 324 tasks are fully balanced… See the full description on the dataset page: https://huggingface.co/datasets/BIM-Edit/BIM-Edit-Tasks.MMH3_Image_Edit_WorkflowThis is just an example of using MiniMax H3 as an image editor. The actual workflow that I use requires several custom nodes, some of which are not published, so this one is simply a bare bones demonstration.
This uses the hybrid MiniMax H3 model from here: https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/tree/main
It uses the custom VAE from here: https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE/tree/main
It uses the LoRA from here:… See the full description on the dataset page: https://huggingface.co/datasets/fizzlepoof/MMH3_Image_Edit_Workflow.ImagePulseV2-Edit-Pose
ImagePulseV2 Dataset - Pose Adjustment
The ImagePulseV2 dataset is a custom-built dataset created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Pose.OralGPT-X-EditImagePulseV2-Edit-Style
ImagePulseV2 Dataset - Style Transfer
The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Style.ImagePulseV2-Edit-HumanFace
ImagePulseV2 Dataset - Facial Expression Editing
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It contains multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-HumanFace.foi-source-catalogue
FOI preservation index
Verified snapshot
Manifest SHA-256: a4ec82ea96d033fc7c196faddf4277c36aac963ca83684904a72247c0c98d486.
The snapshot was downloaded anonymously and validated before this index moved. Read its manifest and coverage report for scope. Public visibility is not country completeness. Source rights apply separately; this card grants no blanket licence to source payloads.
corpus-nz-hathi
Historical New Zealand Parliamentary Debates from HathiTrust
Registry status
Registry ID: edithatogo/corpus-nz-hathi
Family: hathitrust-nz
Repository role: derived_corpus
Canonical dataset: edithatogo/nz-hansard-corpus
Operational status: active
Rights status: rights_vary_by_volume
Authoritative catalog: edithatogo/dataset-estate-registry
Origin and provenance
Origin repository: https://github.com/edithatogo/hathi-nz
Upstream source: HathiTrust… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/corpus-nz-hathi.EditEval
EditEval: The Instruction-Based Benchmark for Text Improvements
This dataset contains the EditEval benchmark data, converted to JSONL and organized by task/dataset.
Subsets
Config
Task
Examples
jfleg
Fluency
1,501
asset
Simplification
2,359
turk
Simplification
2,359
iterater
Mixed (all tasks)
621
iterater_fluency
Fluency
203
iterater_clarity
Clarity
342
iterater_coherence
Coherence
76
stsb_multi_mt
Paraphrasing
153
wnc
Neutralization
1,700… See the full description on the dataset page: https://huggingface.co/datasets/bzz2/EditEval.synthetic-commit-msg-edits
✍️ Commit Message Edits Dataset - 🤖Synthetic
This dataset is a synthetic extension of our expert-labeled commit message edits dataset presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings.
You can check Synthetic tab in our visualization app to browse through the datapoints!
Dataset Structure
Default
Default split contains the synthetic messages generated from expert-labeled dataset by an LLM.… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/synthetic-commit-msg-edits.
