datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.WMGStereo
What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching? (WMGStereo)
Paper | GitHub
WMGStereo is a procedural dataset generator specifically optimized for zero-shot stereo matching performance. This repository contains the WMGStereo-150k dataset, a large-scale synthetic training dataset featuring indoor, nature, and dense "flying" scenes.
Dataset Download
You can download the dataset using the huggingface-cli:
pip install huggingface-cli
huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/princeton-vl/WMGStereo.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.Stereo4D_vlbm
Stereo4D (converted to VLBM format)
This dataset contains 4,687 sequences from the Stereo4D dataset converted to the VLBM-compatible format using preprocess_stereo4d.py. The sequences have been compressed into .tar.gz archives in chunks of 50 sequences per archive.
Scale
Metric
Value
Total sequences
4,687
Image resolution
512 x 512 px
Depth type
Sparse (projected from tracked 3D points)
Dataset Structure
Each sequence directory follows this… See the full description on the dataset page: https://huggingface.co/datasets/ZhengGuangze/Stereo4D_vlbm.raw_primitive_datasets
Datacard
This is the official fine-tuning dataset provided by VLABench (raw data), with 500 episodes each task. The current version includes 10 primitive tasks.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
Download all archive files and use the following command to extract:
cat vlabench_primitive.tar.gz.* | tar -xzvf -
In the resulting… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/raw_primitive_datasets.CALVIN-3D_PCD-ABC_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABC_D.CALVIN-3D_PCD-ABCD_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABCD_D.SynplerHuman
Humans Dataset
Synthetic human image dataset generated for:
On the Role of Visual Realism in Synthetic Data for Estimating 3D Human Pose and Shape
Paper: TODOCode:
Structure
shards/
<dataset_name>/
shard-*.tar
labels/
<dataset_name>_combined.npz
manifest.csv
manifest.csv maps each shard directory to its matching label file.
Labels
Each .npz label file contains frame level annotations, including:
imgname: image path/name matching entries… See the full description on the dataset page: https://huggingface.co/datasets/princeton-vl/SynplerHuman.infinigen-articulated
Infinigen-Articulated Assets
Formerly Infinigen-Sim
VerMultiThis repository contains the data presented in LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.
Project page: https://forjadeforest.github.io/LMM-R1-ProjectPage
FloorPlan-VLN-RxRvln_r2r_rxr_hfov90Innovator-VL-Instruct-ScienceVLN-CEKubric_vlbm
Kubric (converted to VLBM format)
This dataset contains 11,000 sequences from the Kubric (TAPVid3D) dataset converted to the VLBM-compatible format using preprocess_kubric.py. The sequences have been compressed into .tar.gz archives in chunks of 50 sequences per archive.
Scale
Metric
Value
Total sequences
11,000
Frames per sequence
24
Image resolution
512 x 512 px
Depth type
Dense (ground truth)
Dataset Structure
Each sequence directory… See the full description on the dataset page: https://huggingface.co/datasets/ZhengGuangze/Kubric_vlbm.GATE-VLAP-datasets
GATE-VLAP Datasets
Grounded Action Trajectory Embeddings with Vision-Language Action Planning
This repository contains preprocessed datasets from the LIBERO benchmark suite in WebDataset TAR format, specifically designed for training vision-language-action models with semantic action segmentation.
Data Format: WebDataset TAR
We provide datasets in WebDataset TAR format for optimal performance:
✅ Fast loading - Efficient streaming during training✅ Easy downloading - Single… See the full description on the dataset page: https://huggingface.co/datasets/gate-institute/GATE-VLAP-datasets.doc_vla_cache_train
NAVSIM navtrain metric cache
NAVSIM navtrain split 的 metric cache,用于 PDM score 计算(AutoVLA 等做 RL/GRPO 训练时的 reward,
或跑 PDMS 评测)。纯 CPU 产物,与模型无关,一次生成可永久复用。
场景数
103,288(train 101,288 + val 2,000)
大小
35 GB(解包后)
生成
navsim/planning/script/run_metric_caching.py,train_test_split=navtrain
覆盖率
对 navtrain 的 101,288 个训练样本 100% 覆盖,缺 0
每个 metric_cache.pkl(lzma 压缩,约 450 KB)含 PDMScorer 判分所需的全部内容:
ego_state、trajectory(PDM 参考轨迹)、observation(各时刻 agent 占用)、… See the full description on the dataset page: https://huggingface.co/datasets/PhoenixHu/doc_vla_cache_train.FloorPlan-VLN-R2RVLM-SFTVLMBench_datasetVLMEval-MovieChat1kVLM-150M
Dataset Card for VLM-150M
VLM-150M is a large-scale image-text dataset that has been recaptioned using an SFT-enhanced Qwen2VL model to enhance the alignment and detail of textual descriptions.
Dataset Sources
Repository: [https://zxwei.site/hqclip/)
Usage Guide
See https://github.com/w1oves/hqclip/blob/main/README.md#dataset-usage-guide.
animetimm-Danbooru-VLMVLM_img2end_newCALVIN-3D_cam-params
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_cam-params.small-publaynet-wds
Small PubLayNet (WebDataset)
This dataset consists in the first WebDataset shards of PubLayNet from http://storage.googleapis.com/nvdata-publaynet
It is mostly used to test the WebDataset integration within the Hugging Face ecosystem.
PointOdyssey_vlbm_old
PointOdyssey to VLBM Format Conversion Report
This document summarizes the process and results of converting the PointOdyssey training dataset to the Visual Lattice Boltzmann Model (VLBM) format.
Conversion Overview
The PointOdyssey train split was converted using a multi-processed Python script (pointodyssey2vlbm.py). The conversion involved transforming source RGB images, 16-bit depth maps, and coordinate annotations into the standardized format used by the VLBM dataset… See the full description on the dataset page: https://huggingface.co/datasets/ZhengGuangze/PointOdyssey_vlbm_old.Stereo4D_vlbm_old
Stereo4D (converted to VLBM format, quality top-5%)
This dataset contains a quality-filtered subset of Stereo4D sequences converted to the VLBM/Flock4D-compatible format using the conversion tool stereo4d2vlbm.py.
Scale of Source and Filtered Dataset
The original Stereo4D dataset contains 98,112 sequences sourced from in-the-wild stereo videos. To focus on high-quality dynamic content, we applied an automated video quality detection pipeline (quality_detect.py) and… See the full description on the dataset page: https://huggingface.co/datasets/ZhengGuangze/Stereo4D_vlbm_old.vlm3r_sample_10kMVS-Synth_vlbm
MVS-Synth to VLBM Format Conversion Report
Overview
This report documents the conversion of the MVS-Synth GTAV_540 dataset to the VLBM format for use with video-based learning models.
Source Data: MVS-Synth GTAV_540
MVS-Synth is a synthetic multi-view stereo dataset generated from the video game GTA V. The GTAV_540 subset contains:
Property
Value
Number of sequences
120
Frames per sequence
100
Total frames
12,000
Image resolution
960 × 540… See the full description on the dataset page: https://huggingface.co/datasets/ZhengGuangze/MVS-Synth_vlbm.
