datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.3D_Visual_Illusion_Depth_Estimation
3D Visual Illusion Depth Estimation Dataset
Dataset Summary
The 3D Visual Illusion Depth Estimation Dataset is designed for research on stereo and monocular depth estimation in 3D visual illusion scenes.It contains left and right stereo images, depth maps estimated from DepthAnything V2, and illusion-region masks.
Dataset Structure
Each sample in the dataset includes:
left: Left-view RGB image
right: Right-view RGB image
depth: Monocularly estimated depth… See the full description on the dataset page: https://huggingface.co/datasets/AdamYao/3D_Visual_Illusion_Depth_Estimation.ophnet_3d3DCoMPaT200
3DCoMPaT200 Dataset
The 3DCoMPaT200 dataset is a comprehensive collection of 3D objects with compositional part annotations. This repository contains various formats and versions of the dataset organized for different use cases.
📁 Directory Structure
2D Folder
Contains train, validation, and test data in tar format for 10 compositions:
Training set
Validation set
Test set
Each file contains 2D representations of the objects with their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/CoMPaT/3DCoMPaT200.CALVIN-3D_PCD-ABC_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABC_D.CALVIN-3D_PCD-ABCD_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABCD_D.3d-wm-atomic-v2
3d-wm-atomic-v2 — Synthetic CAD construction videos
Per-op animated frame sequences for CadQuery construction programs, with
per-frame atomic op labels in continuous raw mm. Designed as training
data for video → action (IDM) and image → next frame
(video-gen) models.
Format version: v2-continuous-mm (2026-05-29 onward).
One clip = one (case, view) pair
data_*/{bNNNN}/train/shard-NNNNNN.tar
└── {case_uid}_v{NN}/
├── 0000.png .. NNNN.png # 256x256 rendered… See the full description on the dataset page: https://huggingface.co/datasets/hz6666/3d-wm-atomic-v2.hypernet-image-to-3d-data
Hypernet Image-to-3D Dataset
Companion dataset to the hypernet and image-to-3d-deepsdf repositories.
Contents
Path
Size
Description
manifest.json
<1 MB
obj_idx ↔ uid ↔ LVIS category mapping for 1000 shapes
views.json
<1 KB
64 Fibonacci-sphere camera poses
watertight/
~21 GB
985 watertight .obj meshes (filename = uid hash)
sdf_samples/
~3 GB
976 .npz files, each with 200K (point, sdf) pairs
multiview/
~1.3 GB
976 .tar files; each tarball contains 64 PNG… See the full description on the dataset page: https://huggingface.co/datasets/bobthebuilderinternational/hypernet-image-to-3d-data.3D-Alpaca
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
Paper | Project Page | Code
A subset from the 3D-Alpaca dataset of ShapeLLM-Omni: a native multimodal LLM for 3D generation and understanding
Junliang Ye*, Zhengyi Wang*, Ruowen Zhao*, Shenghao Xie, Jun Zhu
Recently, the powerful text-to-image capabilities of GPT-4o have led to growing appreciation for native multimodal large language models. However, its multimodal capabilities remain confined to images… See the full description on the dataset page: https://huggingface.co/datasets/yejunliang23/3D-Alpaca.3d-1k-assets
3D-1K Assets
A curated 3D rigid-object asset library with a 10-category taxonomy and 1,000 GLB assets (one per base object class). Generated by a text → image → RGBA → 3D pipeline (FLUX.2-klein → RMBG-2.0 → TRELLIS.2).
This is the base class layer of a planned 1,000 × 1,000 = 1,000,000 instance dataset for embodied AI / robotic manipulation.
Files
file
size
content
glbs.tar
6.5 GB
1000 GLB assets: 0001.glb … 1000.glb
rgba_images.tar
700 MB
1000 RGBA… See the full description on the dataset page: https://huggingface.co/datasets/Dennis0626/3d-1k-assets.3d-wm-atomic-xycanon2D-to-3D-groundingCALVIN-3D_cam-params
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_cam-params.FAST_3D_SAM3d2d-and-3d-frames3d-generation3d-wm-atomic-v2vtk3d-wm-atomic-v1MixDiffusion-3D-Cache
