datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
XLRS-Bench_visual_grounding_en
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_en.avspeech-visual-audio
AVSpeech Video + Audio
This repository is a media-bearing reconstruction of the public AVSpeech
annotations. Each row represents an already-trimmed segment and keeps the
original source-video timing and target-face-center metadata.
Dataset structure
clip_id: identifier derived as
{youtube_id}_{start_sec:.3f}_{end_sec:.3f}.
avspeech_metadata: JSON containing youtube_id, start_sec, end_sec,
x_center, and y_center from the AVSpeech annotation.
video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.trex-visualizer
T-Rex Dataset Visualizer
A browseable subset of the T-Rex dataset — Tactile-Rich Bimanual Dexterous
Manipulation — collected on a bimanual Dexmate Vega-1 robot equipped with
two Sharpa Wave dexterous hands.
This visualizer subset contains 3,838 short trajectory clips drawn from the
full 100-hour T-Rex collection, organized by (verb, object, hand) so you can
quickly inspect coverage across motion primitives and object categories.
For the full dataset (multi-view RGB, robot… See the full description on the dataset page: https://huggingface.co/datasets/Beakerman0101/trex-visualizer.institutional-books-hl-visual-elements
📚 Institutional Books: Harvard Library — Visual Elements
22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset.
22,622,060 visual elements extracted from 983,004 volumes
766,992,447 o200k_base tokens in AI-generated captions
6 high-level classes of visual elements organized in splits
5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation
The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.Visual-CoT
VisCoT Dataset Card
There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance.
To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/deepcs233/Visual-CoT.visual_robust_robocasa_xvisual-jenga-datasets
Visual Jenga Datasets
This directory contains the original datasets for Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting. Visual Jenga is a novel scene understanding task that involves progressively removing objects from a single image one at a time while keeping the rest of the scene stable. This process reveals object dependencies and provides a new way to evaluate grounded scene understanding by systematically exploring which objects can be removed… See the full description on the dataset page: https://huggingface.co/datasets/konpat/visual-jenga-datasets.DRIM-VisualReasonHardThis repository contains the RL training datasets used in the paper Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
Galgame-VisualNovel-Reupload
Galgame VisualNovel Reupload
This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574.
The goal of this reupload is to restructure the data for easier and more efficient use with the datasets library, instead of having to manually extract each archive file and parse json files of the original dataset.
Loading the entire dataset
To load and stream all voice lines from all games combined, simply load the train split. The… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/Galgame-VisualNovel-Reupload.imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.visual-puzzlesvisual-head
🔍 Visual Head Analysis Dataset
"Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach" (CVPR 2025)
📖 Overview
This dataset contains comprehensive attention analysis results from various Large Multimodal Models (LMMs) across multiple vision-language benchmarks. The data enables research into visual attention patterns, attention head behavior, and multimodal interpretability.
🛠️ Associated Tools
The accompanying codebase… See the full description on the dataset page: https://huggingface.co/datasets/jing-bi/visual-head.cyberseceval3-visual-prompt-injection
Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark
Dataset Details
Dataset Description
This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains.
Language(s): English
License: MIT
Dataset Sources
Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.VisualPRM400K-v1.1
VisualPRM400K-v1.1
[📂 GitHub]
[📜 Paper]
[🆕 Blog]
[🤗 model]
[🤗 dataset]
[🤗 benchmark]
NOTE: VisualPRM400K-v1.1 is a new version of VisualPRM400K, which is used to train VisualPRM-8B-v1.1. Compared to the original version, v1.1 includes additional data sources and prompts during rollout sampling to enhance data diversity.
NOTE: To unzip the archive of images, please first run cat images.zip_* > images.zip and then run unzip images.zip.
VisualPRM400K is a dataset comprising… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VisualPRM400K-v1.1.Maritime_Visual_Tracking_Dataset_MVTD
MVTD: Maritime Visual Tracking Dataset
Overview
MVTD (Maritime Visual Tracking Dataset) is a large-scale benchmark dataset designed specifically for single-object visual tracking (VOT) in maritime environments.It addresses challenges unique to maritime scenes: such as water reflections, low-contrast objects, dynamic backgrounds, scale variation, and severe illumination changes—which are not adequately covered by generic tracking datasets.
The dataset contains 182… See the full description on the dataset page: https://huggingface.co/datasets/AhsanBB/Maritime_Visual_Tracking_Dataset_MVTD.VisualWebInstruct
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs.
Please also checkout our more recent verified version at Huggingface.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct.Maritime_Visual_Tracking_Dataset_MVTD
MVTD: Maritime Visual Tracking Dataset
Overview
MVTD (Maritime Visual Tracking Dataset) is a large-scale benchmark dataset designed specifically for single-object visual tracking (VOT) in maritime environments.It addresses challenges unique to maritime scenes: such as water reflections, low-contrast objects, dynamic backgrounds, scale variation, and severe illumination changes—which are not adequately covered by generic tracking datasets.
The dataset contains 182… See the full description on the dataset page: https://huggingface.co/datasets/othmaneirl/Maritime_Visual_Tracking_Dataset_MVTD.Visual_Privacy_Dataset
VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection
Official dataset for the ICML 2026 paper
VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection
🌐 Project Page: https://vpd-100k.github.io/
📄 Paper: https://arxiv.org/abs/2605.10229
Overview
Visual privacy protection has become increasingly important as people continuously share images and live-stream videos online. Existing visual privacy datasets are generally… See the full description on the dataset page: https://huggingface.co/datasets/XiaoyuSunANU/Visual_Privacy_Dataset.VisualWebInstruct-Recall
Introduction
This is the dataset recalled from Google Search from the seed images.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
XLRS-Bench_visual_grounding_zh
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_zh.visual_masked_distracting_metaworld
Visual Masked Distracting Meta-World (ground-truth masks)
Author: Georgios Tsakoumakis
Thesis: Interaction-Masked Latent Action Models for Object-Aware Manipulation under Visual Distractors (MSc, Imperial College London)
Expert Meta-World manipulation trajectories rendered with dynamic video-background
distractors, augmented with ground-truth segmentation masks and pose for the
manipulated object: the agent mask plus two per-frame fields, object_mask and
object_state.
All… See the full description on the dataset page: https://huggingface.co/datasets/tsakman23/visual_masked_distracting_metaworld.VisualGenome_VG_100K_1_and_2
VisualProbe_traine6-visual-ratingsvisual-reasoning-benchmark-results
Visual Reasoning Benchmark Suite v3.3 · 2005 Tasks · 12 Tracks Equal Weight
本版本以用户最新上传的 visual_reasoning_benchmark_suite_v3_修改 为唯一基础版本,不回退、不覆盖用户已经重绘或修改过的既有数据。完整性比对结果:原基础包中 3283 个既有数据文件全部保持字节级不变。
在此基础上新增并整合:
Nonogram(数织)150 题:45 Easy / 60 Medium / 45 Hard;
Tangram(七巧板)150 题:45 Easy / 60 Medium / 45 Hard;
两个任务的一键生成器、统一生成入口、统一评估入口、雷达图和排行榜支持。
最终总规模:2005 题,12 个 Track。
任务与数量
Task
Count
figure_completion
394
spatial_generation
56
maze_beginner
64… See the full description on the dataset page: https://huggingface.co/datasets/songyiren/visual-reasoning-benchmark-results.visualization-chartsVisualOverload
Dataset Card for VisualOverload
This is a FiftyOne dataset with 2,720 samples.
It is a FiftyOne-format conversion of the original
paulgavrikov/visualoverload
dataset (CVPR 2026). All credit for the data, annotations, and benchmark design belongs to
the original authors — please see Citation and
Dataset Sources.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/VisualOverload.3D_Visual_Illusion_Depth_Estimation
3D Visual Illusion Depth Estimation Dataset
Dataset Summary
The 3D Visual Illusion Depth Estimation Dataset is designed for research on stereo and monocular depth estimation in 3D visual illusion scenes.It contains left and right stereo images, depth maps estimated from DepthAnything V2, and illusion-region masks.
Dataset Structure
Each sample in the dataset includes:
left: Left-view RGB image
right: Right-view RGB image
depth: Monocularly estimated depth… See the full description on the dataset page: https://huggingface.co/datasets/AdamYao/3D_Visual_Illusion_Depth_Estimation.visual_ai_at_neurips2025
Dataset Card for neurips-2025-vision-papers
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/visual_ai_at_neurips2025")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/visual_ai_at_neurips2025.Graph200K
VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning
[Paper] [Project Page] [Github]
[🤗 Online Demo]
[🤗 Full Model Card (Diffusers)] [🤗 LoRA Model Card (Diffusers)]
Graph200k is a large-scale dataset containing a wide range of distinct tasks of image generation. If you find Graph200k is helpful, please consider to star ⭐ the Github Repo. Thanks!
📰 News
[2025-5-15] 🤗🤗🤗 VisualCloze has been merged into the… See the full description on the dataset page: https://huggingface.co/datasets/VisualCloze/Graph200K.
