datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-edit-simplerswesmith-cleanHQ-Edit
Dataset Card for HQ-EDIT
HQ-Edit, a high-quality instruction-based image editing dataset with total 197,350 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3.
HQ-Edit’s high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/HQ-Edit.MathCanvas-Edit
MathCanvas-Edit Dataset
🚀 Data Usage
from datasets import load_dataset
dataset = load_dataset("shiwk24/MathCanvas-Edit")
print(dataset)
📖 Overview
MathCanvas-Edit is a large-scale dataset containing 5.2 million step-by-step editing trajectories, forming a crucial component of the [MathCanvas] framework. MathCanvas is designed to endow Unified Large Multimodal Models (LMMs) with intrinsic… See the full description on the dataset page: https://huggingface.co/datasets/shiwk24/MathCanvas-Edit.NHR-Edit
NoHumanRequired (NHR) Dataset for image editing
🌐 NHR Website |
📜 NHR Paper on arXiv |
💻 GitHub Repository |
🤗 NHR-Edit Dataset (part2) |
🤗 BAGEL-NHR-Edit |
❗️ Important: This is Part 1 of the Dataset ❗️
Please be aware that this repository contains the first part of the full NHR-Edit dataset. To have the complete training data, you must also download the second part.
➡️ Click Here to Access Part 2 on… See the full description on the dataset page: https://huggingface.co/datasets/iitolstykh/NHR-Edit.edition_2408_jxcai-scale-hle-public-questions-readymade
edition_2408_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2408_jxcai-scale-hle-public-questions-readymade.edition_2618_jxcai-scale-hle-public-questions-readymade
edition_2618_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2618_jxcai-scale-hle-public-questions-readymade.edition_2071_jxcai-scale-hle-public-questions-readymade
edition_2071_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2071_jxcai-scale-hle-public-questions-readymade.edition_2443_jxcai-scale-hle-public-questions-readymade
edition_2443_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2443_jxcai-scale-hle-public-questions-readymade.multi_reference_image_editing
Multi-Reference Instruction-Based Image Editing Dataset
Overview
This dataset contains 20,000 high-resolution image pairs and multi-modal instructions designed for training advanced image-to-image editing models. It combines two complementary example types: 10,000 reference-grounded edits, where structural or stylistic changes are driven by up to three provided visual reference images, and 10,000 occlusion-based inpainting/outpainting edits, where the model must… See the full description on the dataset page: https://huggingface.co/datasets/molbal/multi_reference_image_editing.edition_2419_jxcai-scale-hle-public-questions-readymade
edition_2419_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2419_jxcai-scale-hle-public-questions-readymade.edition_2637_jxcai-scale-hle-public-questions-readymade
edition_2637_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2637_jxcai-scale-hle-public-questions-readymade.edition_2791_jxcai-scale-hle-public-questions-readymade
edition_2791_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2791_jxcai-scale-hle-public-questions-readymade.SweSmith-RL-Datasetedition_3125_jxcai-scale-hle-public-questions-readymade
edition_3125_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_3125_jxcai-scale-hle-public-questions-readymade.ramanv-image-editing-pairsedition_2899_jxcai-scale-hle-public-questions-readymade
edition_2899_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2899_jxcai-scale-hle-public-questions-readymade.fyi-archive-nz
fyi-archive-nz
Derived FOI-O candidate layer
Any separately versioned FOI-O/NLP output is an inferred candidate layer for
research evaluation. It is not a gold label set, a certified legal finding,
legal advice, or a replacement for the immutable raw archive. Candidate
publication requires a separately approved dataset identity and target.
Dataset card for incremental, read-only preservation of
fyi.org.nz — the New Zealand Official Information Act
(OIA) request… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/fyi-archive-nz.edition_2074_jxcai-scale-hle-public-questions-readymade
edition_2074_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2074_jxcai-scale-hle-public-questions-readymade.kiwi_edit_training_data
RefVIE (Kiwi-Edit Training Data)
Project Page | Paper | GitHub
RefVIE is a large-scale dataset tailored for instruction-reference-following video editing tasks, introduced in the paper "Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance".
The dataset was constructed using a scalable data generation pipeline that transforms existing video editing pairs into high-fidelity training quadruplets. It leverages image generative models to create synthesized reference… See the full description on the dataset page: https://huggingface.co/datasets/linyq/kiwi_edit_training_data.r2e-dockers-rllm-v1ramanv-image-real-style-editorialedition_2602_jxcai-scale-hle-public-questions-readymade
edition_2602_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2602_jxcai-scale-hle-public-questions-readymade.edition_2186_jxcai-scale-hle-public-questions-readymade
edition_2186_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2186_jxcai-scale-hle-public-questions-readymade.editbench
EditBench Dataset
This dataset contains code editing tasks extracted from the EditBench evaluation framework specifically designed for evaluating model performance on code editing tasks. It is provided as a test-only benchmark. Each sample includes:
Please check out https://github.com/waynchi/editbench for our full evaluation harness.
Core Files (Python)
original_code.py: Starting code file
highlighted_code.py: Specific section of code to be modified
instruction.txt:… See the full description on the dataset page: https://huggingface.co/datasets/copilot-arena/editbench.emu_edit_test_set
Dataset Card for the Emu Edit Test Set
Dataset Summary
To create a benchmark for image editing we first define seven different categories of potential image editing operations: background alteration (background), comprehensive image changes (global), style alteration (style), object removal (remove), object addition (add), localized modifications (local), and color/texture alterations (texture).
Then, we utilize the diverse set of input images from the MagicBrush… See the full description on the dataset page: https://huggingface.co/datasets/facebook/emu_edit_test_set.NHR-Edit-part2
NoHumanRequired (NHR) Dataset for image editing
🌐 NHR Website |
📜 NHR Paper on arXiv |
💻 GitHub Repository |
🤗 NHR-Edit Dataset (part1) |
🤗 BAGEL-NHR-Edit-V2 |
❗️ Important: This is Part 2 of the Dataset ❗️
Please be aware that this repository contains the second part of the full NHR-Edit dataset. To have the complete training data, you must also download the first part.
➡️ Click Here to Access Part 1 on… See the full description on the dataset page: https://huggingface.co/datasets/iitolstykh/NHR-Edit-part2.editlens_iclredition_0558_ryanmarten-OpenThoughts-1k-sample-readymade
edition_0558_ryanmarten-OpenThoughts-1k-sample-readymade
A Readymade by TheFactoryX
Original Dataset
ryanmarten/OpenThoughts-1k-sample
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0558_ryanmarten-OpenThoughts-1k-sample-readymade.edition_2590_jxcai-scale-hle-public-questions-readymade
edition_2590_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2590_jxcai-scale-hle-public-questions-readymade.
