datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hy-Embodied-0.5-VLA-Data
Hy-Embodied-0.5-VLA
From Vision-Language-Action Models to a Real-World Robot Learning Stack
Tencent Robotics X × Tencent Hy Team
📖 Abstract
We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.Innovator-VL-Instruct-46M
Innovator-VL-Instruct-46M
Paper | Code
🤗🤗 The data is being uploaded continuously
Introduction
To further enhance the model’s ability to handle a broad range of visual tasks with accurate, grounded, and instruction-aligned responses, we perform full-parameter visual instruction supervised fine-tuning (SFT).This SFT stage serves as a critical bridge between multimodal pretraining and subsequent reinforcement learning, providing both general capability coverage and a… See the full description on the dataset page: https://huggingface.co/datasets/InnovatorLab/Innovator-VL-Instruct-46M.picbreeder-vlm-archive
Picbreeder-VLM Archive
Every image evolved by the swarm of vision-language-model "breeders" in
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
(GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the
lineage graphs, and the analysis artifacts behind the paper and blog.
The original Picbreeder (Secretan et al., 2008) let crowds of
humans collaboratively evolve images from
CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.VLNVerse_sceneVLADBenchLavalObjaverseDataset
Laval Objaverse Dataset
vLAR Group | SIGGRAPH Asia 2026
A large-scale, high-quality dataset for multi-view relighting.
📖 Dataset Summary
The Laval Objaverse Dataset is a comprehensive dataset designed for multi-view relighting and novel view synthesis tasks. It combines high-quality 3D assets from Objaverse with realistic, diverse illumination conditions… See the full description on the dataset page: https://huggingface.co/datasets/vLAR/LavalObjaverseDataset.PhysInOnePhysInOne: Visual Physics Learning and Reasoning in One Suite
vLAR Group |
The Hong Kong Polytechnic University |
Syai Singapore |
Meta
CVPR 2026
🧭 Navigation
📌 Summary
🚀 Release Timetable
📦 Repositories & Downloads
📊 Data Splits
🧱 3D Assets
🛠️ Data Processing
🏆 Leaderboard Evaluation Data
🎞️ Rendered Data
1. Download Scripts
2. Install Dependencies… See the full description on the dataset page: https://huggingface.co/datasets/vLAR/PhysInOne.vlabench_primitive_pretrain_lerobot
Datacard
This is the official VLABench primitive pretraining dataset converted to the
LeRobot format. The dataset contains language-conditioned manipulation
trajectories collected with a Franka Panda robot in VLABench simulation.
This LeRobot version is hosted at:
https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code:… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot.VLNVerse_datavllm-control-arena
vLLM Main Tasks Dataset
AI coding tasks generated from vLLM git commits
Dataset Description
This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work.
Dataset Structure
The dataset contains the following columns:
commit_hash: The git commit hash
parent_hash: The parent commit hash
commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.vlm-teacher-embeddingmikasa-robo-vla-rldsInFlux-Synth
InFlux++ Synth
InFlux++ Synth is a large-scale synthetic training dataset for the InFlux project, providing per-frame ground truth camera intrinsics and camera pose for videos with dynamic intrinsics.
The dataset contains 441,840 annotated frames from 1,841 procedurally generated high-resolution videos. Every video contains 240 frames at a resolution of 1280 × 720. The dataset spans indoor and nature scenes and features changing zoom and focus, dynamic objects, and realistic… See the full description on the dataset page: https://huggingface.co/datasets/princeton-vl/InFlux-Synth.LayeredFlow-Syn
LayeredFlow-Syn Extracted Ground Truth
This repository contains extracted LayeredFlow ground-truth annotations in Parquet format.
Layout
data/<scene>/<sample>.parquet
For example:
data/0/0_0.parquet
Each Parquet file stores one extracted sample. Rows correspond to files from the extracted
sample directory and include a leftmost visualization image preview when available,
relative_path, num_bytes, sha256, and binary content.
Rows are ordered by frame0/left… See the full description on the dataset page: https://huggingface.co/datasets/princeton-vl/LayeredFlow-Syn.mv3dpt-datasets
Multi-View 3D Point Tracking Datasets
This repository hosts the training and evaluation datasets associated with the paper Multi-View 3D Point Tracking.
Project Page: https://ethz-vlg.github.io/mvtracker/
Code/Github Repository: https://github.com/ethz-vlg/mvtracker
Abstract
We introduce the first data-driven multi-view 3D point tracker, designed to track arbitrary points in dynamic scenes using multiple camera views. Unlike existing monocular trackers, which struggle… See the full description on the dataset page: https://huggingface.co/datasets/ethz-vlg/mv3dpt-datasets.eval-resultsVLN_Dataset_2This repository contains encrypted visual features for an ongoing academic research project. Decryption keys are managed internally for reproducibility.
PaddleOCR-VL_demogpt-edit-simplervlabench-assetsBench2Drive-VL-base
Bench2Drive-VL: Full-Stack Software for Closed-Loop Autonomous Driving with Vision Language Models
Project Page | GitHub | Paper
Bench2Drive-VL is a comprehensive closed-loop benchmark for Vision-Language Models in Autonomous Driving (VLM4AD). It extends the Bench2Drive benchmark by introducing closed-loop evaluation and the DriveCommenter expert model for automated annotation.
This repository contains the natural language annotations for the Bench2Drive-Base1000 dataset. These… See the full description on the dataset page: https://huggingface.co/datasets/Telkwevr/Bench2Drive-VL-base.Recap-DataComp-1B
Dataset Card for Recap-DataComp-1B
Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions.
Dataset Details
Dataset Description
Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM.
Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.vlm_test_imagesBunch of random test cases for vision language in the wild.
Nemotron-VLM-Dataset-v2
Nemotron-VLM-Dataset v2
Versions
Date
Commit
Changes
2025-11-05
head
Fix nights_cot dataset. Fix/filter broken <think> entries. Update fintabnet instructions. Update indexes.
2025-10-28
214051e
Initial Release
Dataset Description
Following up on Llama Nemotron VLM Dataset V1 with 3 million samples, we are releasing the Nemotron VLM Dataset V2 with almost three times as many high-quality samples.
This time, our focus was on three… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-VLM-Dataset-v2.OpenVLMRecords
OpenVLM Records
Here we maintain all the evaluation records generated by VLMEvalKit,
which also reflects on the OpenVLM Leaderboard.
Before using the scripts to browse and utilize those record files, you should first have VLMEvalKit installed
(use pip install -e . --no-deps when you encounter some dependency errors).
Naming System & Record Browsing
In this repo, records are organized with the following naming system:
The record file of evaluating MLLM VLM-A on the… See the full description on the dataset page: https://huggingface.co/datasets/VLMEval/OpenVLMRecords.collected_demos_trainingLET-KUAVO-VLA-1.0-Dataset
LET-KUAVO-VLA-1.0-Dataset
VLAC-Cut-FullData
VLAC-Cut-FullData
VLAC-Cut-FullData is the full-data release for VLAC-Cut. It provides the complete raw-data archive set, benchmark-style JSON files, and a lightweight frame-extraction workflow for reproducing evaluation on the released benchmark protocol.
Contents
benchmark_style_all/
train/video_progress_benchmark_file.json
test_expert_seen/video_progress_benchmark_file.json
test_expert_unseen/video_progress_benchmark_file.json… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/VLAC-Cut-FullData.vlm_resultsGPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.
