datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vstat
VSTAT: Visual State Tracking Benchmark
VSTAT is a video-based benchmark for evaluating the visual state tracking
capability of Multimodal Large Language Models (MLLMs). It contains 834 video
clips paired with 1,500 questions whose answers cannot be inferred from any
single keyframe or short segment.
Dataset Composition
Split
Videos
Questions
synthetic
450
550
self_recorded
80
100
youtube
304
850
Total
834
1,500
Files… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/vstat.Vision2Web
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
[🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results]
Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle.
This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.miracl-vision
MIRACL-VISION
MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark.
This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.prof_report__SG161222-Realistic_Vision_V1.4__multi__24
Dataset Card for "prof_report__SG161222-Realistic_Vision_V1.4__multi__24"
More Information needed
DisciplineGen-1MEgoHaFL
EgoHaFL: Egocentric 3D Hand Forecasting Dataset with Language Instruction
EgoHaFL is a dataset designed for egocentric (first-person) 3D hand forecasting with accompanying natural language instructions.
It was introduced in the paper SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting.
The dataset contains short video clips, text descriptions, camera intrinsics, and detailed MANO-based 3D hand annotations.
The dataset supports research in 3D hand… See the full description on the dataset page: https://huggingface.co/datasets/ut-vision/EgoHaFL.PVL-BA-Bench
PVL-BA-Bench
PVL-BA-Bench is a large-scale bundle adjustment benchmark dataset from Polar-vision Lab. It provides public release metadata, interactive browser viewers, and downloadable PVL-BA, COLMAP, and BAL packages for photogrammetric optimization research.
Release Links
Static release index and interactive viewers: https://pub-2c28bdf6e62548919c47727a9b969dda.r2.dev/index.html
Downloadable packages in this dataset repository: packages/{pvl-ba,colmap… See the full description on the dataset page: https://huggingface.co/datasets/Polar-vision/PVL-BA-Bench.Art-Vision-Question-Answering-Dataset
Art Vision Question Answering Dataset
🎨 A curated dataset for training AI models on digital artwork analysis and visual question answering.
Dataset Overview
This dataset contains 577 question-answer pairs extracted from artwork conversations, designed for training multimodal AI models on art analysis tasks.
✨ Key Features
🖼️ Visual Thumbnails: Artwork images displayed directly in the dataset viewer
💬 Rich Q&A: Expert-level questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/OneEyeDJ/Art-Vision-Question-Answering-Dataset.rollout_groot_vision_only_Relational_PlacementThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_vision_only_Relational_Placement.trade_vision_dataset
TradeVision: Hierarchical Physical Business & Multimodal Retail Provenance Dataset
This dataset is continuously seeded from OpenStreetMap, matched to Google Place IDs, harvested for temporal store photos, and enriched with zero-shot computer vision using Hugging Face Hub native pipelines.
Dataset Structure
The dataset is partitioned into three relational subsets loadable via Hugging Face datasets:
from datasets import load_dataset
# 1. Load Canonical Businesses… See the full description on the dataset page: https://huggingface.co/datasets/drksci/trade_vision_dataset.touch-vision-v3_20260902_152740This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/nchristo-synaptics/touch-vision-v3_20260902_152740.touch-vision-v1_20260902_110913This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/nchristo-synaptics/touch-vision-v1_20260902_110913.touch-vision-v2_20260902_122712This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/nchristo-synaptics/touch-vision-v2_20260902_122712.eval_act_vision_charger_50epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_tactile_follower",
"total_episodes": 10,
"total_frames": 8969,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aryankakad/eval_act_vision_charger_50ep.drawer_vision_only_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 75,
"total_frames": 25156,
"total_tasks": 1,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:75"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LSY-lab/drawer_vision_only_v2.multimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.rollout_groot_vision_only_Referential_DisambiguationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_vision_only_Referential_Disambiguation.stack_vision_only_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 11455,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LSY-lab/stack_vision_only_v2.muse-k2-vision-pilot-20260910
Muse → K2 bridge: first training experiment
Prepared September 10, 2026. This experiment tests whether training a connector
lets the frozen IFM/K2-Horizon-7B decoder use the existing Muse-Glimmer visual
encoder. It does not retrain the vision encoder or K2, and it does not establish
general screenshot, document, natural-image, or visual reasoning capability.
Authorized budget and selected first hardware
The user authorized an initial inexpensive Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/txgsync/muse-k2-vision-pilot-20260910.rollout_groot_vision_only_Counting_20260831_180259This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_vision_only_Counting_20260831_180259.InfiniBench_challlenge
Download
git lfs install
git clone https://huggingface.co/datasets/Vision-CAIR/InfiniBench_challlenge
cd InfiniBench_challlenge
rm -rf .git/
How to Decompress the Videos
To extract the videos from the compressed files, follow these steps:
Open a terminal and navigate to the videos directory:
cd videos
Combine all the split archive parts into a single .tar.gz file:
cat test_videos.tar.gz.part_* > test_videos.tar.gz
Extract the contents of the archive:
tar -xvf… See the full description on the dataset page: https://huggingface.co/datasets/Vision-CAIR/InfiniBench_challlenge.rollout_groot_vision_only_State_RecognitionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_vision_only_State_Recognition.ditflow_drawer_vision_only_v1_evalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 26,
"total_frames": 11562,
"total_tasks": 1,
"total_videos": 78,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:26"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/danielsanjosepro/ditflow_drawer_vision_only_v1_eval.lm-eval-results-Nitral-AI-Eris_PrimeV3.05-Vision-7B-private
Dataset Card for Evaluation run of Nitral-AI/Eris_PrimeV3.05-Vision-7B
Dataset automatically created during the evaluation run of model Nitral-AI/Eris_PrimeV3.05-Vision-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Nitral-AI-Eris_PrimeV3.05-Vision-7B-private.eval_smolvla_vision_charger_20epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_tactile_follower",
"total_episodes": 10,
"total_frames": 7498,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aryankakad/eval_smolvla_vision_charger_20ep.eval_smolvla_vision_charger_50epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_tactile_follower",
"total_episodes": 10,
"total_frames": 8224,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aryankakad/eval_smolvla_vision_charger_50ep.africa-synth-disability-visual-impairment-low-vision-all
Visual Impairment & Low Vision Services (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-disability-visual-impairment-low-vision-all.rollout_groot_vision_only_Relational_Placement_20260830_143726This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_vision_only_Relational_Placement_20260830_143726.rollout_groot_vision_only_Referential_Disambiguation_20260829_160104This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_vision_only_Referential_Disambiguation_20260829_160104.rollout_groot_vision_only_Size_Recognition_20260829_123751This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_vision_only_Size_Recognition_20260829_123751.
