datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpic
GPIC: A Giant Permissive Image Corpus for Visual Generation
Keshigeyan Chandrasegaran*1,
Kyle Sargent*1,
Suchir Agarwal1,
Michael Jang1,
Michael Poli1,2,
Juan Carlos Niebles1,4,
Justin Johnson3,
Jiajun Wu1,
Li Fei-Fei1
1 Stanford University
2 Radical Numerics
3 University of Michigan
4 Salesforce… See the full description on the dataset page: https://huggingface.co/datasets/stanford-vision-lab/gpic.course-assetsEgoBrain
[ICLR 2026] EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
東京大学 The Univerisity of Tokyo X 微軟亞洲研究院 Microsoft Research Asia
Nie Lin
·
Yansen Wang
·
Dongqi Han
·
Weibang Jiang
·
Jingyuan Li
·
Ryosuke Furuta
·
Yoichi Sato*
·
Dongsheng Li*
·
*(Co-corresponding authors)*
This is the official dataset repository of our ICLR 2026 paper "EgoBrain: Synergizing Minds and Eyes For Human Action… See the full description on the dataset page: https://huggingface.co/datasets/ut-vision/EgoBrain.VSI-590K
VSI-590K
Website | Paper | GitHub | Models
Authors: Shusheng Yang*, Jihan Yang*, Pinzhi Huang†, Ellis Brown†, et al.
VSI-590K is a large-scale spatially-focused instruction-tuning dataset focusing on spatial reasoning. The dataset is curated from diverse sources and carefully annotated.
Quick Start
import json
# Load from JSONL file
with open('vsi_590k.jsonl', 'r') as f:
for line in f:
sample = json.loads(line.strip())
print(sample)
break… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-590K.SenseNova-Vision-Corpus-50M
Vision as Unified Multimodal Generation
English | 简体中文
This repository contains the dataset for the paper Vision as Unified Multimodal Generation.
SenseNova Vision Corpus 50M
Overview
SenseNova Vision Corpus 50M (SN-VC-50M) is a large-scale multimodal vision corpus designed for unified training across diverse visual understanding and geometry-oriented tasks. The dataset is curated to address a common limitation of existing… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/SenseNova-Vision-Corpus-50M.Cambrian-10M
Cambrian-10M Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-10M is a comprehensive dataset designed for instruction tuning, particularly in multimodal settings involving visual interaction data. The dataset is crafted to address the scarcity of high-quality multimodal instruction-tuning data and to maintain the language abilities of multimodal large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-10M.Cambrian-S-3M
Cambrian-S-3M
TLDR: This is a collection of open-source video instruction tuning data used in Cambrian-S's third training stage.
Overview
Cambrian-S-3M combines three video instruction datasets:
Cambrian-S-3M
LLaVA-Video-178K
LLaVA-Hound (ShareGPTVideo)
Prerequisites
Hugging Face CLI: pip install -U "huggingface_hub[cli]==0.36.0"
Sufficient disk space (~5 TB recommended)
hf command should be available after installing huggingface_hub
Setup… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-S-3M.pipeline-cctv-analyticskaz-vision-50kVSI-Bench
Dataset
arXiv
Website
Code
VSI-Bench
VSI-Bench-Debiased v1
[!IMPORTANT]
[Aug. 9, 2026] PROVENANCE UPDATE: The existing "Debiased" subset is VSI-Bench-Debiased v1, a designer-in-the-loop manual pilot created with bespoke per-question-type filtering heuristics. It predates and was not generated by the automated Iterative Bias Pruning (IBP) algorithm. We retain v1 for reproducibility and will version any future automated subset separately.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-Bench.open-vision-banana-snvc-train-full
SNVC-50M v5_full — Multi-Task Vision Dataset
Description
This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs).
Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.CV-Bench
Cambrian Vision-Centric Benchmark (CV-Bench)
This repository contains the Cambrian Vision-Centric Benchmark (CV-Bench), introduced in Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.
Files
The test*.parquet files contain the dataset annotations and images pre-loaded for processing with HF Datasets.
These can be loaded in 3 different configurations using… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/CV-Bench.Cambrian-Alignment
Cambrian-Alignment Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V.
Getting Started with Cambrian Alignment Data
Before you start, ensure you have sufficient storage space to download and process the data.
Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.VisionArena-Chat
VisionArena-Battle: 30K Real-World Image Conversations with Pairwise Preference Votes
200k single and multi-turn chats between users and VLM's collected on Chatbot Arena.
WARNING: Images may contain inappropriate content.
Dataset Details
200K conversations
45 VLM's
138 languages
~43k unique images
Question Category Tags (Captioning, OCR, Entity Recognition, Coding, Homework, Diagram, Humor, Creative Writing, Refusal)
Dataset Description
200,000… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/VisionArena-Chat.OpenMath-Vision-CoT-10kvision-flan_191-task_1k
🚀 Vision-Flan Dataset
vision-flan_191-task-1k is a human-labeled visual instruction tuning dataset consisting of 191 diverse tasks and 1,000 examples for each task.
It is constructed for visual instruction tuning and for building large-scale vision-language models.
Paper or blog for more information:
https://github.com/VT-NLP/MultiInstruct/
https://vision-flan.github.io/
Paper coming soon 😊
Citation
Paper coming soon 😊. If you use Vision-Flan, please use the… See the full description on the dataset page: https://huggingface.co/datasets/Vision-Flan/vision-flan_191-task_1k.ui-vision
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Introduction
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/ui-vision.prof_images_blip__SG161222-Realistic_Vision_V1.4
Dataset Card for "prof_images_blip__SG161222-Realistic_Vision_V1.4"
More Information needed
Vision-OPD-6K
Vision-OPD-6K: Training Data for Vision-OPD
Overview
Vision-OPD proposes a regional-to-global self-distillation framework that transfers the model's own privileged regional perception to its full-image policy, without external teacher models, ground-truth labels, reward verifiers, or inference-time tool use.
Vision-OPD instantiates two conditional policies from the same MLLM:
A crop-conditioned teacher that observes the evidence-centered crop as a privileged… See the full description on the dataset page: https://huggingface.co/datasets/yuanqianhao/Vision-OPD-6K.vision-feedback-mix-binarized
Dataset Card for Vision-Feedback-Mix-Binarized
Introduction
This dataset aims to provide large-scale vision feedback data.
It is a combination of the following high-quality vision feedback datasets:
zhiqings/LLaVA-Human-Preference-10K: 9,422 samples
MMInstruction/VLFeedback: 80,258 samples
YiyangAiLab/POVID_preference_data_for_VLLMs: 17,184 samples
openbmb/RLHF-V-Dataset: 5,733 samples
openbmb/RLAIF-V-Dataset: 83,132 samples
We also offer a cleaned version in… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/vision-feedback-mix-binarized.pisa-experiments
Pisa Experiments
This repository contains the PisaBench, training data, model checkpoints, introduced in PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop.
PisaBench
Real World Videos
We curate a dataset comprising 361 videos demonstrating the dropping task.Each video begins with an object suspended by an invisible wire in the first frame. We cut the video clips to begin as soon as the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/pisa-experiments.vstat
VSTAT: Visual State Tracking Benchmark
VSTAT is a video-based benchmark for evaluating the visual state tracking
capability of Multimodal Large Language Models (MLLMs). It contains 834 video
clips paired with 1,500 questions whose answers cannot be inferred from any
single keyframe or short segment.
Dataset Composition
Split
Videos
Questions
synthetic
450
550
self_recorded
80
100
youtube
304
850
Total
834
1,500
Files… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/vstat.VisionUnlearningEvaluationTestbedsWake-Vision
Dataset Card for Wake Vision
Dataset Description
"Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of
current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally,
it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age,
subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.vision-adapter-embeddingsvision-arena-bench-v0.1
VisionArena-Bench: An automatic eval pipeline to estimate model preference rankings
An automatic benchmark of 500 diverse user prompts that can be used to cheaply approximate Chatbot Arena model rankings via automatic benchmarking with VLM as a judge.
Dataset Sources
Repository: https://github.com/lm-sys/FastChat
Paper: https://arxiv.org/abs/2412.08687
Automatic Evaluation Code: Coming Soon!
Dataset Structure
question_id: The unique hash representing the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/vision-arena-bench-v0.1.Websight_Mantis_Datachest-xray-pneumoniaDataset Summary
The dataset is organized into 3 folders (train, test, val) and contains subfolders for each image category (Pneumonia/Normal). There are 5,863 X-Ray images (JPEG) and 2 categories (Pneumonia/Normal).
Chest X-ray images (anterior-posterior) were selected from retrospective cohorts of pediatric patients of one to five years old from Guangzhou Women and Children’s Medical Center, Guangzhou. All chest X-ray imaging was performed as part of patients’ routine clinical care.
For… See the full description on the dataset page: https://huggingface.co/datasets/hf-vision/chest-xray-pneumonia.Cambrian10M_For_MantisVision2Web
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
[🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results]
Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle.
This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.
