datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ScaleEdit-12M
ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework
📌 Overview
The largest open-source instruction-based image editing dataset to date.
ScaleEdit-12M contains 11.8 million rigorously verified instruction–image pairs spanning 23 task families across diverse real and synthetic visual domains. It was constructed using ScaleEditor, a fully open-source hierarchical multi-agent framework that eliminates… See the full description on the dataset page: https://huggingface.co/datasets/InternVL-U/ScaleEdit-12M.scaleedit-filtered-6m
ScaleEdit Filtered 6M
This repository contains a portable selection manifest for high-quality samples
from ScaleEdit-12M. It does not redistribute the source images. Download the
original ScaleEdit-12M Parquet files separately, then join each manifest row to
the source file named by source_relative_path at zero-based row_index.
Selection
For each evaluated source row, the latest successful stage-2 review was used.
A row is included when result.final_decision ==… See the full description on the dataset page: https://huggingface.co/datasets/QingyuShi/scaleedit-filtered-6m.
