datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
institutional-books-hl-visual-elements
📚 Institutional Books: Harvard Library — Visual Elements
22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset.
22,622,060 visual elements extracted from 983,004 volumes
766,992,447 o200k_base tokens in AI-generated captions
6 high-level classes of visual elements organized in splits
5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation
The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.VisualWebInstruct
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs.
Please also checkout our more recent verified version at Huggingface.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct.VisualWebInstruct-Recall
Introduction
This is the dataset recalled from Google Search from the seed images.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
visual_masked_distracting_metaworld
Visual Masked Distracting Meta-World (ground-truth masks)
Author: Georgios Tsakoumakis
Thesis: Interaction-Masked Latent Action Models for Object-Aware Manipulation under Visual Distractors (MSc, Imperial College London)
Expert Meta-World manipulation trajectories rendered with dynamic video-background
distractors, augmented with ground-truth segmentation masks and pose for the
manipulated object: the agent mask plus two per-frame fields, object_mask and
object_state.
All… See the full description on the dataset page: https://huggingface.co/datasets/tsakman23/visual_masked_distracting_metaworld.VisualPuzzles
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
🏠 Homepage | 📊 VisualPuzzles | 💻 Github | 📄 Arxiv | 📕 PDF | 🖥️ Zeno Model Output
Overview
VisualPuzzles is a multimodal benchmark specifically designed to evaluate reasoning abilitiesin large models while deliberately minimizing reliance on domain-specific knowledge.
Key features:
1168 diverse puzzles
5 reasoning categories: Algorithmic, Analogical, Deductive, Inductive, Spatial… See the full description on the dataset page: https://huggingface.co/datasets/neulab/VisualPuzzles.visual_distracting_metaworld
Visual Distracting Meta-World with Agent Masks
Visual Distracting Meta-World with Agent Masks contains successful expert
trajectories for all 50 Meta-World
MT50 v3 manipulation tasks. Every visual observation includes agent masks: a
ground-truth robot-arm mask obtained from the simulator, a SAM 2.1
mask predicted from the clean observation, and a separate SAM 2.1 mask predicted
from the distracted observation. Each step pairs these masks with the same
underlying simulator state… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/visual_distracting_metaworld.VisualWebInstruct-Seed
Introduction
This is the seed dataset we used to conduct Google Search.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
VisualSphinx-V1-Raw
🦁 VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
VisualSphinx is the largest fully-synthetic open-source dataset providing vision logic puzzles. It consists of over 660K automatically generated logical visual puzzles. Each logical puzzle is grounded with an interpretable rule and accompanied by both correct answers and plausible distractors.
🌐 Project Website - Learn more about VisualSphinx
📖 Technical Report - Discover the methodology and technical details… See the full description on the dataset page: https://huggingface.co/datasets/VisualSphinx/VisualSphinx-V1-Raw.Visual-Searchvisual_masked_distracting_metaworld_sam
Visual Masked Distracting Meta-World (ground-truth + SAM masks)
Author: Georgios Tsakoumakis
Thesis: Interaction-Masked Latent Action Models for Object-Aware Manipulation under Visual Distractors (MSc, Imperial College London)
Expert Meta-World manipulation trajectories rendered with dynamic video-background
distractors, carrying both ground-truth and SAM-predicted segmentation masks for the
agent and the manipulated object, plus the manipulated object's pose.
This is the SAM… See the full description on the dataset page: https://huggingface.co/datasets/tsakman23/visual_masked_distracting_metaworld_sam.oxford-iiit-pet-vl-enriched
Visualize on Visual Layer
Oxford-IIIT-Pets-VL-Enriched
An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues!
With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues help to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.visual_26k_reasoningVisualWebBench
VisualWebBench
Dataset for the paper: VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
🌐 Homepage | 🐍 GitHub | 📖 arXiv
Introduction
We introduce VisualWebBench, a multimodal benchmark designed to assess the understanding and grounding capabilities of MLLMs in web scenarios. VisualWebBench consists of seven tasks, and comprises 1.5K human-curated instances from 139 real websites, covering 87 sub-domains. We evaluate 14… See the full description on the dataset page: https://huggingface.co/datasets/visualwebbench/VisualWebBench.VisualToolBench
VisToolBench Dataset
A benchmark dataset for evaluating vision-language models on tool-use tasks.
Dataset Statistics
Total samples: 1204
Single-turn: 603
Multi-turn: 601
Schema
Column
Type
Description
id
string
Unique task identifier
turncase
string
Either "single-turn" or "multi-turn"
num_turns
int
Number of conversation turns (1 for single-turn)
prompt_category
string
Task category (e.g., "medical", "scientific", "general")
eval_focus… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/VisualToolBench.visual-qa-llama-format
Open Paws Visual Qa Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Multimodal Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.visual_distracting_metaworld_with_masksvisual_distracting_control_suite
Visual Distracting Control Suite Benchmark
This dataset contains expert trajectories generated by a Proximal Policy Optimization (PPO) reinforcement learning agent trained on 4 environments of the Distracting Control Suite. For each environment we collect data with different levels of distraction, which we define below, and masks for the agent.
Levels of distraction:
None: Vanilla DeepMind Control Suite without visual distractions. The environment uses the default static background… See the full description on the dataset page: https://huggingface.co/datasets/hamza-adnan/visual_distracting_control_suite.spatial-visual-reasoning-66kdocument-visual-retrieval-test
Model Card: Document Visual Retrieval Test (internal)
Dataset Overview
This dataset is designed to evaluate the performance of visual retrievers by testing their ability to match a query to a relevant image. Each of the three examples in this dataset contains a text query and an associated image, which is a scanned page from the foundational "Attention is All You Need" paper. The purpose of this dataset is to facilitate the evaluation of visual retrievers, where the… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/document-visual-retrieval-test.v1
Knowledge in Visual Synthesis
This dataset contains prompt–image examples for evaluating and studying
knowledge-intensive visual synthesis. Samples are organized by contributor as
dataset subsets (configs), with each upload version exposed as a split.
Dataset structure
Subset
Splits
byx
v1, v2
yuner
v1, v2
zanyi
v1, v2, v3
jiayu
v1, v2, v3
sherry
v1, v2
yujunz
v1
The byx/v1 split contains 140 unique prompts and 300 generated images. For… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.coco-2014-vl-enriched
Visualize on Visual Layer
COCO-2014-VL-Enriched
An enriched version of the COCO 2014 dataset with label issues! The label issues help to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original image filename from the COCO dataset.
image: Image data in the form of PIL Image.
label_bbox: Bounding box annotations from the COCO dataset. Consists of bounding box coordinates, confidence scores, and labels… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/coco-2014-vl-enriched.Visual-JigsawVisual-TableQA
🧠 Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images
Welcome to Visual-TableQA, a project designed to generate high-quality synthetic question-answer datasets associated to images of tables. This resource is ideal for training and evaluating models on visually-grounded table understanding tasks such as document QA, table parsing, and multimodal reasoning.
🚀 Latest Update
We have refreshed the dataset with newly generated QA pairs created by… See the full description on the dataset page: https://huggingface.co/datasets/AI-4-Everyone/Visual-TableQA.visual_genomeOcean_R1_collected_visual_dataSOC-Training-Data-Visualization
Paper Link
SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding
Code repo
Code for Generation
Citation
@misc{huang2025sossyntheticobjectsegments,
title={SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding},
author={Weikai Huang and Jieyu Zhang and Taoyang Jia and Chenhao Zheng and Ziqi Gao and Jae Sung Park and Ranjay Krishna},
year={2025},
eprint={2510.09110},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/SOC-Training-Data-Visualization.VisualSphinx-V1-Raw-PanelsARO-Visual-RelationPersonaChat-Qwen-Image-2512-enhanced
