datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3DCode
Project page
Paper
Code
3dcodebench.com
arXiv:2606.01057
gaoypeng/3dcodebench
News
[06/01/2026] Paper released on arXiv: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code.
Note. This is an open-source reproduction of 3DCodeBench.
⚠️ Under final check. The 3DCodeData/ code is still undergoing final
quality review and may contain occasional issues (non-executable scripts, mismatched
captions/renders, or imperfect geometry). If you run… See the full description on the dataset page: https://huggingface.co/datasets/YipengGao/3DCode.PanoCity
PanoCity Dataset
A Large-Scale Aerial Panoramic Dataset for 3D Scene Understanding
📊 Dataset Statistics
Attribute
Value
Total Size
1.4 TB
Cities
Beijing (20 blocks), Jinan (76 blocks), Ningbo (41 blocks)
Panoramic RGB Images
119,537 (2048×4096)
Panoramic Depth Maps
119,537 (2048×4096)
Total Images
239,074
Image Format
PNG
📂 Data Structure
PanoCity/
├── splits_config.json # Official train/test splits
├──… See the full description on the dataset page: https://huggingface.co/datasets/YijingGuo/PanoCity.VOST-TAS
[NeurIPS 2025] Tracking and Understanding Object Transformations
If you like our project, please give us a star ⭐ on GitHub for the latest update.
💡 Description
Dataset Visualizations: GitHub
Paper: arXiv:2511.04678
Project Page: tubelet-graph.github.io
Project Repository: GitHub
Point of Contact: Yihong Sun
📊 Dataset Overview
VOST-TAS (TrackAnyState) is an extended version of the VOST validation set with explicit transformation annotations for tracking and… See the full description on the dataset page: https://huggingface.co/datasets/yihongs/VOST-TAS.RefRef_additionalRefRef: A Synthetic Dataset and Benchmark for Reconstructing Refractive and Reflective Objects
Yue Yin ·
Enze Tao ·
Weijian Deng ·
Dylan Campbell
About
This repository provides additional data for the RefRef dataset.
Citation
@misc{yin2025refrefsyntheticdatasetbenchmark,
title={RefRef: A Synthetic Dataset and Benchmark for Reconstructing Refractive and Reflective Objects},
author={Yue Yin and Enze Tao and… See the full description on the dataset page: https://huggingface.co/datasets/yinyue27/RefRef_additional.robomme_preprocessed_data
RoboMME Training Data (Pickle Format)
Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code
This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments.
.
├── data # zipped pickle files
├── features # zipped precompute siglip embeddings
├── meta # statistics for robomme
├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.robomme_data_lerobot
RoboMME Training Data (LeRobot Format)
Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code
This repo contains Lerobot format of RoboMME training data.
We use this in our Diffusion Policy training code
libero_cotThis dataset was created using LeRobot.
It contains embodied Chain-of-Thought (CoT) demonstrations for the LIBERO benchmark, featuring paired reasoning and action traces. It was curated as part of the DeepThinkVLA project using a two-stage data engine that distills key frames with a cloud LVLM and scales to full trajectories via a fine-tuned local VLM.
Paper: DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
Repository: https://github.com/OpenBMB/DeepThinkVLA… See the full description on the dataset page: https://huggingface.co/datasets/yinchenghust/libero_cot.copyblobtest-doc-assetsrl-game-traces-death-stranding-2
死亡搁浅2
This public dataset repository contains gameplay trace data uploaded from F:\死亡搁浅2.
Contents
Files: 521
Total local size: 479.86 GB
Generated: 2026-06-12T04:47:09+00:00
File Types
.jsonl: 160
.png: 123
.json: 120
.mkv: 40
.txt: 40
.parquet: 38
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage, audio, and asset… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-death-stranding-2.OPV2V-H
OPV2V-H dataset
Based on the original OPV2V dataset, we supplemented 16-line, 32-line lidar data for each agent, as well as 4 depth cameras. Annotations will be shared with the original OPV2V, so please download the original OPV2V dataset as well.
Depth Data
OPV2V-H-depth.zip stores the depth camera data. You can directly uncompress them:
unzip OPV2V-H-depth.zip
LiDAR Data
OPV2V-H-LiDAR-partxx is a volume-compressed slice. They store 16-line and 32-line LiDAR data. You can… See the full description on the dataset page: https://huggingface.co/datasets/yifanlu/OPV2V-H.SARLANG-1M
SARLANG-1M
SARLANG-1M is a large-scale benchmark tailored for multimodal SAR image understanding, with a primary focus on integrating SAR with textual modality. SARLANG-1M comprises more than 1 million high-quality SAR image-text pairs collected from over 59 cities worldwide. It features hierarchical resolutions (ranging from 0.1 to 25 meters), fine-grained semantic descriptions (including both concise and detailed captions), diverse remote sensing categories (1,696… See the full description on the dataset page: https://huggingface.co/datasets/YiminJimmy/SARLANG-1M.rl-game-traces-rise-of-the-tomb-raider
古墓丽影:崛起
This public dataset repository contains gameplay trace data uploaded from F:\古墓丽影:崛起.
Contents
Files: 647
Total local size: 496.11 GB
Generated: 2026-06-06T01:03:39+00:00
File Types
.jsonl: 196
.json: 149
.png: 129
.parquet: 49
.mkv: 49
.txt: 49
.jpg: 15
.exe: 11
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-rise-of-the-tomb-raider.MIRA
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
Dataset Description
MIRA (Multimodal Imagination for Reasoning Assessment) evaluates whether MLLMs can think while drawing—i.e., generate and use intermediate visual representations (sketches, diagrams, trajectories) as part of reasoning.MIRA includes 546 carefully curated problems spanning 20 task types across four domains:
Euclidean Geometry (EG)
Physics-Based Reasoning (PBR)… See the full description on the dataset page: https://huggingface.co/datasets/YiyangAiLab/MIRA.controlnet-testingGaussianEditor_Resultrl-game-traces-horizon-forbidden-west
地平线之西之绝境
This public dataset repository contains gameplay trace data uploaded from F:\地平线之西之绝境.
Contents
Files: 442
Total local size: 375.37 GB
Generated: 2026-06-07T09:25:23+00:00
File Types
.jsonl: 132
.json: 100
.png: 79
.parquet: 33
.mkv: 33
.txt: 33
.jpg: 32
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage, audio… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-horizon-forbidden-west.MMS-VPR
MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments
Overview
MMS-VPR is the first large-scale multimodal street-level visual place recognition dataset featuring comprehensive integration of images, videos, and rich textual annotations with day–night coverage and a 7-year temporal span in dense pedestrian-only environments.
MMS-VPR comprises 110,529 images and 2,527 video clips… See the full description on the dataset page: https://huggingface.co/datasets/Yiwei-Ou/MMS-VPR.rl-game-traces-resident-evil-4-remake
生化危机4重制版
This public dataset repository contains gameplay trace data uploaded from F:\生化危机4重制版.
Contents
Files: 199
Total local size: 131.17 GB
Generated: 2026-06-13T19:51:14+00:00
File Types
.jsonl: 59
.png: 48
.json: 47
.parquet: 15
.mkv: 15
.txt: 15
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage, audio, and asset… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-resident-evil-4-remake.I2V-CompBench
I2V-CompBench
A compositional image-to-video (I2V) generation benchmark spanning 7 evaluation dimensions, with first-frame images derived from TIP-I2V and refined text prompts produced by a dual VLM/LLM pipeline.
⚠️ License: CC BY-NC 4.0 (inherits from TIP-I2V). Non-commercial use only.
📦 Versions
This repository hosts two parallel snapshots of the same benchmark. Pick the layout that fits your tooling.
Version
Path
Questions
Layout
Best for
v2 ⭐… See the full description on the dataset page: https://huggingface.co/datasets/YiningZ2002/I2V-CompBench.blogimgRoboTwin_instruct-pix2pix
数据集介绍
Contributor: Shenghao Yang, Yimou Wu
数据简介
本数据集是通过https://github.com/TianxingChen/RoboTwin [1]单臂机器人模拟器在 block_hammer_beat, block_handover, blocks_stack_easy 三个任务上间隔50帧采样得到的,其数据规模如下:
block_hammer_beat block_handover blocks_stack_easy
train 100 100 100
val 10 10 10
test 10 10 10
整体数据规模如下:
train val test
num 300 30 30… See the full description on the dataset page: https://huggingface.co/datasets/YimouWu/RoboTwin_instruct-pix2pix.Urban-ImageNet
🏙️ Urban-ImageNet
A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception from Social Media Imagery.
Urban-ImageNet fills a critical gap between computer vision and urban studies by treating cities not simply as visual scenes, but as lived, socially produced, and experientially activated spaces.
Overview
ImageNet taught models to recognise objects. Urban-ImageNet teaches them to understand how people experience cities.… See the full description on the dataset page: https://huggingface.co/datasets/Yiwei-Ou/Urban-ImageNet.MEHSI
MEHSI
Multi-Exposure real HSI denoising (MEHSI) dataset for "Real Noise Decoupling for Hyperspectral Image Denoising" accepted at AAAI 2026.
Details
We aim to capture a larger and more varied noise level dataset, so we control exposure time to 1/20, 1/50, and 1/100 of the reference image captured. We utilize an SOC710-VP hyperspectral camera to capture HSIs.
The collected clean HSIs are averaged, and the paired data are manually aligned and calibrated. The collected data… See the full description on the dataset page: https://huggingface.co/datasets/YingkaiZhang/MEHSI.RLAIF-V-Dataset
Dataset Card for RLAIF-V-Dataset
This dataset was introduced in RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness.
GitHub
This dataset was also used in MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
News:
[2025.09.18] 🎉 Our data is used in the powerful MiniCPM-V 4.5 model, which represents a state-of-the-art end-side MLLM achieving GPT-4o level performance!
[2025.03.01] 🎉 RLAIF-V is accepted by CVPR 2025!… See the full description on the dataset page: https://huggingface.co/datasets/YigeLi/RLAIF-V-Dataset.libero-100This is the processed dataset for LIBERO 90 in huggingface LeRobot format to be used for π₀ fine-tuning. The original OpenVLA repo here does not contain LIBERO 100, for RLDS -> LeRobot conversion, so I'm publishing this here. LIBERO 90 data has been preprocessed to remove no-ops, unsuccessful trajectories, and had its image observations flipped back upright according to the script OpenVLA authors provided.
rl-game-traces-euro-truck-simulator-2
欧洲卡车模拟2
This public dataset repository contains gameplay trace data uploaded from F:\欧洲卡车模拟2.
Contents
Files: 471
Total local size: 344.36 GB
Generated: 2026-06-11T05:43:41+00:00
File Types
.jsonl: 144
.json: 111
.png: 108
.parquet: 36
.mkv: 36
.txt: 36
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage, audio, and asset… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-euro-truck-simulator-2.UniMM-Chat
Dataset Card for UniMM-Chat
Dataset Summary
UniMM-Chat dataset is an open-source, knowledge-intensive, and multi-round multimodal dialogue data powered by GPT-3.5, which consists of 1.1M diverse instructions.
UniMM-Chat leverages complementary annotations from different VL datasets and employs GPT-3.5 to generate multi-turn dialogues corresponding to each image, resulting in 117,238 dialogues, with an average of 9.89 turns per dialogue.
A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/Yirany/UniMM-Chat.rl-game-traces-civilization-6
文明6
This public dataset repository contains gameplay trace data uploaded from F:\文明6.
Contents
Files: 413
Total local size: 363.61 GB
Generated: 2026-06-09T09:17:06+00:00
File Types
.jsonl: 123
.json: 106
.png: 90
.parquet: 32
.txt: 31
.mkv: 30
.xlsx: 1
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage, audio, and asset… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-civilization-6.UnsafeBench
Dataset Card for Dataset Name
[Update]: we added the caption/prompt information (if there is one) in case other researchers need it. It is not used in our study though.
The dataset consists of 10K safe/unsafe images of 11 different types of unsafe content and two sources (real-world VS AI-generated).
Dataset Details
Source
# Safe
# Unsafe
# All
LAION-5B (real-world)
3,228
1,832
5,060
Lexica (AI-generated)
2,870
2,216
5,086
All
6,098
4,048
10,146… See the full description on the dataset page: https://huggingface.co/datasets/yiting/UnsafeBench.
