datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MagicBrush
Dataset Card for MagicBrush
Dataset Summary
MagicBrush is the first large-scale, manually-annotated instruction-guided image editing dataset covering diverse scenarios single-turn, multi-turn, mask-provided, and mask-free editing. MagicBrush comprises 10K (source image, instruction, target image) triples, which is sufficient to train large-scale image editing models.
Please check our website to explore more visual results.
Dataset Structure
"img_id" (str):… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/MagicBrush.Multimodal-Mind2Web
Dataset Summary
Multimodal-Mind2Web is the multimodal version of Mind2Web, a dataset for developing and evaluating generalist agents
for the web that can follow language instructions to complete complex tasks on any website. In this dataset, we align each HTML document in the dataset with
its corresponding webpage screenshot image from the Mind2Web raw dump. This multimodal version addresses the inconvenience of loading images from the ~300GB Mind2Web Raw Dump.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/Multimodal-Mind2Web.bioscan-traits
Dataset Card for BIOSCAN-Traits
Dataset Details
Dataset Description
BIOSCAN-Traits is a trait-level annotation dataset for fine-grained insect imagery. Derived from BIOSCAN-5M, it provides morphology-centric natural language trait descriptions automatically generated by a two-stage pipeline: (1) a Sparse Autoencoder (SAE) trained on DINOv2 visual features identifies species-level salient visual parts (wings, legs, antennae, etc.), and (2) a Multimodal LLM… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/bioscan-traits.AutoElicit-Exec
AutoElicit-Exec Dataset
Project Page | Paper | GitHub
AutoElicit-Exec is a human-verified dataset of 132 execution trajectories exhibiting unintended behaviors from typical benign execution. All trajectories are elicited from frontier CUAs (i.e., Claude 4.5 Haiku and Claude 4.5 Opus) using AutoElicit, which perturbs benign instructions from OSWorld to increase the likelihood of unintended harm while keeping instructions realistic and benign. This dataset is designed to provide… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/AutoElicit-Exec.GUI-Drag-dataset
GUI-Drag Dataset
Project Page | Paper | GitHub
GUI-Drag is a diverse dataset of 161K text dragging examples synthesized through a scalable pipeline. It is designed to advance Graphical User Interface (GUI) grounding beyond simple clicking, focusing on the essential interaction of dragging the mouse to select and manipulate textual content.
Our dataset is built upon the Uground, Jedi, and additional public paper-style academic document screenshots.
NOTE: Before you use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/GUI-Drag-dataset.osuGossipcoposutwitter-fatiao_siwaFQ-2025.07.25-1948705972240945547-zhD_OsuORwBZjIoS-part1
