datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
course-assetsSenseNova-Vision-Corpus-50M
Vision as Unified Multimodal Generation
English | 简体中文
This repository contains the dataset for the paper Vision as Unified Multimodal Generation.
SenseNova Vision Corpus 50M
Overview
SenseNova Vision Corpus 50M (SN-VC-50M) is a large-scale multimodal vision corpus designed for unified training across diverse visual understanding and geometry-oriented tasks. The dataset is curated to address a common limitation of existing… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/SenseNova-Vision-Corpus-50M.kaz-vision-50kCV-Bench
Cambrian Vision-Centric Benchmark (CV-Bench)
This repository contains the Cambrian Vision-Centric Benchmark (CV-Bench), introduced in Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.
Files
The test*.parquet files contain the dataset annotations and images pre-loaded for processing with HF Datasets.
These can be loaded in 3 different configurations using… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/CV-Bench.open-vision-banana-snvc-train-full
SNVC-50M v5_full — Multi-Task Vision Dataset
Description
This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs).
Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.Cambrian-Alignment
Cambrian-Alignment Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V.
Getting Started with Cambrian Alignment Data
Before you start, ensure you have sufficient storage space to download and process the data.
Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.VisionArena-Chat
VisionArena-Battle: 30K Real-World Image Conversations with Pairwise Preference Votes
200k single and multi-turn chats between users and VLM's collected on Chatbot Arena.
WARNING: Images may contain inappropriate content.
Dataset Details
200K conversations
45 VLM's
138 languages
~43k unique images
Question Category Tags (Captioning, OCR, Entity Recognition, Coding, Homework, Diagram, Humor, Creative Writing, Refusal)
Dataset Description
200,000… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/VisionArena-Chat.OpenMath-Vision-CoT-10kvision-flan_191-task_1k
🚀 Vision-Flan Dataset
vision-flan_191-task-1k is a human-labeled visual instruction tuning dataset consisting of 191 diverse tasks and 1,000 examples for each task.
It is constructed for visual instruction tuning and for building large-scale vision-language models.
Paper or blog for more information:
https://github.com/VT-NLP/MultiInstruct/
https://vision-flan.github.io/
Paper coming soon 😊
Citation
Paper coming soon 😊. If you use Vision-Flan, please use the… See the full description on the dataset page: https://huggingface.co/datasets/Vision-Flan/vision-flan_191-task_1k.ui-vision
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Introduction
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/ui-vision.prof_images_blip__SG161222-Realistic_Vision_V1.4
Dataset Card for "prof_images_blip__SG161222-Realistic_Vision_V1.4"
More Information needed
vision-feedback-mix-binarized
Dataset Card for Vision-Feedback-Mix-Binarized
Introduction
This dataset aims to provide large-scale vision feedback data.
It is a combination of the following high-quality vision feedback datasets:
zhiqings/LLaVA-Human-Preference-10K: 9,422 samples
MMInstruction/VLFeedback: 80,258 samples
YiyangAiLab/POVID_preference_data_for_VLLMs: 17,184 samples
openbmb/RLHF-V-Dataset: 5,733 samples
openbmb/RLAIF-V-Dataset: 83,132 samples
We also offer a cleaned version in… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/vision-feedback-mix-binarized.Wake-Vision
Dataset Card for Wake Vision
Dataset Description
"Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of
current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally,
it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age,
subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.chest-xray-pneumoniaDataset Summary
The dataset is organized into 3 folders (train, test, val) and contains subfolders for each image category (Pneumonia/Normal). There are 5,863 X-Ray images (JPEG) and 2 categories (Pneumonia/Normal).
Chest X-ray images (anterior-posterior) were selected from retrospective cohorts of pediatric patients of one to five years old from Guangzhou Women and Children’s Medical Center, Guangzhou. All chest X-ray imaging was performed as part of patients’ routine clinical care.
For… See the full description on the dataset page: https://huggingface.co/datasets/hf-vision/chest-xray-pneumonia.Vision2Web
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
[🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results]
Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle.
This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.vision-arena-bench-v0.1
VisionArena-Bench: An automatic eval pipeline to estimate model preference rankings
An automatic benchmark of 500 diverse user prompts that can be used to cheaply approximate Chatbot Arena model rankings via automatic benchmarking with VLM as a judge.
Dataset Sources
Repository: https://github.com/lm-sys/FastChat
Paper: https://arxiv.org/abs/2412.08687
Automatic Evaluation Code: Coming Soon!
Dataset Structure
question_id: The unique hash representing the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/vision-arena-bench-v0.1.lllab-vision-banana-datasetmiracl-vision
MIRACL-VISION
MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark.
This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.VisionThink-Smart-Train
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
Senqiao/VisionThink-Smart-Train
This is the training dataset used for our Efficient Reasoning VLM on general VQA tasks.VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning [Paper]
Senqiao Yang,
Junyi Li,
Xin Lai,
Bei Yu,
Hengshuang Zhao,
Jiaya Jia
Highlights
Our VisionThink leverages reinforcement learning to autonomously learn whether… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/VisionThink-Smart-Train.retrico-vision-benchmarkGRADEVision-SR1-47KVisionThink-General-Train
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
Senqiao/VisionThink-General-Train
This is the training dataset used for our Reasoning VLM on general VQA tasks.
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning [Paper]
Senqiao Yang,
Junyi Li,
Xin Lai,
Bei Yu,
Hengshuang Zhao,
Jiaya Jia
Highlights
Our VisionThink leverages reinforcement learning to autonomously learn whether to… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/VisionThink-General-Train.InfiniBench
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
[ Project Page] [📝 arXiv Paper] [🤗 Download] [🏆Leaderboard]
🔥 News
[2025-08-14] 🏆 2025 ICCV CLVL - Long Video Understanding Challenge (InfiniBench) is now live! Submit your predictions for test set evaluation.
[2025-06-10] This is a new released version of Infinibench.
Overview:
InfiniBench skill set comprising eight skills. The right side represents skill… See the full description on the dataset page: https://huggingface.co/datasets/Vision-CAIR/InfiniBench.FIRM-GenV-ICAL
V-ICAL AI Configuration Dataset
This private dataset contains the AI configurations and visual demonstrations used by V-ICAL for interactive play, batch evaluation, ablation studies, and trajectory auditing.
Layout
<game-id>/<configuration-name>/config.json
<game-id>/<configuration-name>/frames/*.jpg
<game-id>/<configuration-name>/video.mp4
Some configurations also reference preload action sequences maintained by the V-ICAL project.
Download
From… See the full description on the dataset page: https://huggingface.co/datasets/VisionXLab/V-ICAL.MWS-Vision-Bench
MWS-Vision-Bench
🇷🇺 Русскоязычное описание ниже / Russian summary below.
MWS Vision Bench — the first Russian-language business-OCR benchmark designed for multimodal large language models (MLLMs).This is the validation split - publicly available for open evaluation and comparison.🧩 Paper is coming soon.
Language Configs
This dataset provides three Hugging Face configs for the same benchmark split:
default — Russian questions
en — English questions
zh —… See the full description on the dataset page: https://huggingface.co/datasets/MTSAIR/MWS-Vision-Bench.PaperLens-Vision
PaperLens-Vision
Vision version of the OpenReview-ICLR and arXiv PaperLens datasets.
Each row is one unique paper. We release all extracted papers — not every paper here is used in our downstream training/eval sets. The papers that are used are denoted by the references field, which lists every internal (release_name, release_split) pair the paper belongs to (a single paper can belong to multiple). reconstruction.py reads this field to materialize the original sharegpt data.json… See the full description on the dataset page: https://huggingface.co/datasets/skonan/PaperLens-Vision.MTA-Vision-DeepSearchvision-opd-vqa14k-fullimage-curriculum-v8
Vision-OPD VQA14K Full-Image Curriculum v8
Private single-image visual-question-answering dataset.
Split
Rows
Train
14,000
Diagnostic validation
609
The repository contains 14,609 content-addressed media files (4,580,273,467 bytes). Paths in both Parquet files are relative to the repository root and follow media/<sha256-prefix>/<filename>.
from pathlib import Path
import pyarrow.parquet as pq
from huggingface_hub import snapshot_download
root =… See the full description on the dataset page: https://huggingface.co/datasets/yyy051007/vision-opd-vqa14k-fullimage-curriculum-v8.
