datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3DCode
Project page
Paper
Code
3dcodebench.com
arXiv:2606.01057
gaoypeng/3dcodebench
News
[06/01/2026] Paper released on arXiv: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code.
Note. This is an open-source reproduction of 3DCodeBench.
⚠️ Under final check. The 3DCodeData/ code is still undergoing final
quality review and may contain occasional issues (non-executable scripts, mismatched
captions/renders, or imperfect geometry). If you run… See the full description on the dataset page: https://huggingface.co/datasets/YipengGao/3DCode.3D-Front3d-front-rgb3D-ADAMRepository for the 3D-ADAM (3D Anomaly Detection in Additive Manufacturing) Dataset. This is the raw data for our complete dataset, separated by part-instance to allow users to utilise the dataset as desired.
We provide a single-camera (using the MechMind-Nano) subset prepared for unsupervised training at anomaly detection, localisation and segementation tasks through the anomalib library in a separate repository: here
Our ArXiv paper can also be found here: 3D-ADAM Dataset
This project has… See the full description on the dataset page: https://huggingface.co/datasets/pmchard/3D-ADAM.robotwin_3d
RoboTwin 2.0 — 3D (RGB + Depth)
Bimanual manipulation data from the RoboTwin 2.0 simulator, in LeRobot v2.1 format, with per-camera ground-truth depth alongside RGB.
Tasks
50
Episodes
27,500 (550 per task, contiguous)
Frames
6,183,813
Robot
ALOHA-style bimanual, 14-DoF
Control rate
50 Hz
Cameras
3 (cam_high, cam_left_wrist, cam_right_wrist)
Resolution
240 × 320
Language instructions
1,039,891 unique corpus-wide; 100 entries per episode
Total size… See the full description on the dataset page: https://huggingface.co/datasets/flex-pi/robotwin_3d.3dfront_render_views3d-arenaFor more information, visit the 3D Arena Space.
Inputs are sourced from iso3D.
To assist with easily running inputs, are input image URLs are provided in inputs.txt.
3d-front-araria_synthetic_envs_mcmc_3dgs_newLicense Notice:This dataset is derived from the Aria Dataset.It follows the Aria Synthetic Environments Dataset License Agreement.See Aria License for details.
DepR-3D-FRONT
Dataset for DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion
Project Page
|
arXiv
File Structure
File
Optional
Description
pickled_data
Raw data (images, etc.) from InstPIFu
instpifu_mask
Instance masks from InstPIFu
metadata
JSONL metadata for scenes
panoptic
Panoptic segmentation maps we rendered
depth
✅
Estimated depth with Depth Pro
grounded_sam
✅
Estimated segmentation with Grounded SAM… See the full description on the dataset page: https://huggingface.co/datasets/zx1239856/DepR-3D-FRONT.3d-spatial-reasoning-23d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.3D-ZeFThis is the dataset page for 3D-ZeF, the first RGB 3D multiple object tracking dataset of its kind. Zebrafish is a widely used model organism for studying neurological disorders, social anxiety, and more. Behavioral analysis can be a critical part of such research and it has traditionally been conducted manually, which is an expensive, time consuming, and subjective task.
Dataset Setup
We have used an off-the-shelf setup for capturing the dataset, which consists of two GoPro cameras… See the full description on the dataset page: https://huggingface.co/datasets/vapaau/3D-ZeF.3d-spatial-reasoning-13D-Front
3D-Front (MIDI-3D)
Github | Project Page | Paper | Original Dataset
1. Dataset Introduction
TL;DR: This dataset processes 3D-Front into organized 3d scenes paired with rendered multi-view images and surfaces, which are used in MIDI-3D. Each scene contains:
3D models (.glb)
Point cloud (.npy)
Rendered multi-view images in RGB, depth, normal, with camera information
2. Data Extraction
sudo apt-get install git-lfs
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/huanngzh/3D-Front.SAGE-3D_VLN_Data
SAGE-3D VLN Data: Vision-Language Navigation Dataset with Hierarchical Instructions
Paper | Project Page | Code
A comprehensive VLN dataset featuring 2 million trajectory-instruction pairs across 1,000 indoor scenes, with hierarchical instruction design covering high-level semantic goals to low-level control commands.
Overview of SAGE-3D VLN Data. SAGE-3D VLN Data includes a hierarchical instruction, and two major task types (VLN + No-goal).
📢 News… See the full description on the dataset page: https://huggingface.co/datasets/spatialverse/SAGE-3D_VLN_Data.SAGE-3D_Collision_Mesh
SAGE-3D Collision Mesh: Physics-Enabled Collision Bodies for 3D Gaussian Scenes
Paper | Project Page | Code
High-precision collision geometry dataset extracted from 1,000 indoor Mesh scenes, enabling physically accurate navigation and interaction in virtual environments.
Collision Mesh of InteriorGS data captured on Issac Sim 5.0.
📢 News
2025-12-15: Released SAGE-3D Collision Mesh dataset with collision bodies for 1000 InteriorGS scenes.… See the full description on the dataset page: https://huggingface.co/datasets/spatialverse/SAGE-3D_Collision_Mesh.3dfront-render-views3d-defectbench
3D-DefectBench
A benchmark and controlled study of vision-language models (VLMs) as judges for
fine-grained defect detection in text-to-3D generated assets. Each asset carries
a 9-dimensional binary defect vector over five geometry and four texture defect
categories, three of which are prompt-conditioned.
Version 1.1 adds the complete set of 1,000 benchmark GLB assets, the cell-level
VLM prediction table, and the TRELLIS cross-generator prompts and prediction
tables. TRELLIS… See the full description on the dataset page: https://huggingface.co/datasets/aieval2026/3d-defectbench.3D_native_rot_gen3dgs3d-pinn-bedrock3d-spatial-reasonings2oThis repo contains the data for S2O: Static to Openable Enhancement for Articulated 3D Objects.
See the code on GitHub and the paper for details. Please cite S2O [1] if you use ACD.
We provide the mesh, point cloud, and metadata for the two datasets used in S2O.
PM-Openable - This is a subset of 648 openable objects from full PartNet-Mobility [2]. We use a train/val/test split of 460/95/93 objects.
Articulated Container Dataset (ACD) [1] - We take openable container objects from HSSD [3]… See the full description on the dataset page: https://huggingface.co/datasets/3dlg-hcvc/s2o.3D-FUTURE3dfront-render-diffuse3D-dungeon-crawler-video-v2
3D Dungeon Crawler Video v2
32,000 deterministic 28-second observational Unity episodes.
The canonical split contains 16,000 pretrain, 14,000 training,
1,000 test, and 1,000 evaluation episodes.
Unity renders at 512x288 for supersampling. Videos are stored at
256x144, 30 fps, H.264. Training samples every third frame,
yielding 280 frames and an 18x32 visual-token grid per episode.
manifest.jsonl is authoritative for asset paths. Each record points to one MP4 and one
NPZ… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-video-v2.4DNeX-10M
4DNeX-10M Dataset
📄 Paper | 🚀 Project Page | 💻 GitHub
Introduction
4DNeX-10M is a large-scale hybrid dataset introduced in the paper "4DNeX: Feed-Forward 4D Generative Modeling Made Easy".
The dataset aggregates monocular videos from diverse sources, including both static and dynamic scenes, accompanied by high-quality pseudo 4D annotations generated using state-of-the-art 3D and 4D reconstruction methods. The dataset enables joint modeling of RGB appearance and… See the full description on the dataset page: https://huggingface.co/datasets/3DTopia/4DNeX-10M.3DSpatialBench
3DSpatialBench
This dataset repository contains a processed CSV file for 3D spatial benchmarking.
File: filtered_processed_.csv
Source path (local): /pfs/gaohongcheng/3ddata/filtered_processed_.csv
Please update this README with schema, column descriptions, and licensing info.
quickstart-3d
Dataset Card for quickstart-3d
This is a FiftyOne dataset with 200 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = fouh.load_from_hub("Voxel51/quickstart-3d")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/quickstart-3d.
