datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ntu-rgbdSyn4D_RGBD
Dataset Card for Syn4D RGBD
Syn4D is a large-scale, fully-synthetic multiview dataset of dynamic scenes designed to advance research in 4D reconstruction, depth estimation, 3D point tracking, novel-view synthesis, and human pose estimation. It provides dense, complete, and accurate geometric annotations — including per-pixel depth maps, multi-view camera trajectories, dense long-range 3D point tracks, and parametric SMPL-X human body annotations — across a diverse collection of… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Syn4D_RGBD.TUM_RGBD-SLAMTUM_RGBD-SLAMOVIS_RGBD
OVIS_RGBD
The dataset is for paper "Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation".
You can find the usages in GitHub.
The original frame images and annotations are from OVIS. We use DepthAnythingV2 to perform monocular depth estimation on all images. We concatenate depth map on the channel demension and each image is in RGBD format.
Citations
@InProceedings{niu2025,
author = {Niu, Quanzhu and Zhou, Yikang and Chen, Shihao and… See the full description on the dataset page: https://huggingface.co/datasets/QuanzhuNiu/OVIS_RGBD.RGB-D-SegmentEgocentricBodiesannotations_creators:
- other
language:
- en
language_creators:
- other
license:
- odc-by
multilinguality:
- monolingual
pretty_name: 'RGB-D-SegmentEgocentricBodies '
size_categories:
- 1K<n<10K
source_datasets:
- original
tags:
- egocentric segmentation
- extended reality
- xr
- human-body
- mixed-reality
- avatar
task_categories:
- image-segmentation
- depth-estimation
task_ids:
- semantic-segmentation
- features:
- name: image
dtype: image
- name: depth
dtype: image… See the full description on the dataset page: https://huggingface.co/datasets/ExtendedRealityLab/RGB-D-SegmentEgocentricBodies.SPAR-7M-RGBD
📦 Spatial Perception And Reasoning Dataset – RGBD (SPAR-7M-RGBD)
A large-scale multimodal dataset for 3D-aware spatial perception and reasoning in vision-language models.
SPAR-7M-RGBD extends the original SPAR-7M with additional depths, camera intrinsics, and pose information. It contains over 7 million QA pairs across 33 spatial tasks, built from 4,500+ richly annotated indoor 3D scenes.
This version supports single-view, multi-view, and… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhango/SPAR-7M-RGBD.egocentric-stereo-rgbd
Hub Egocentric: Stereo RGB-D
9 egocentric stereo RGB-D clips with dense metric depth, IMU, 6-DoF VIO pose, and calibration. Each clip carries a self-contained LeRobot v3.0 dataset (loads on lerobot >= 0.6.0) plus side-by-side stereo, mono, a colorized depth preview, and a Foxglove MCAP recording.
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on StereoLabs ZED X Mini. Egocentric, human-demonstration data (passive; no robot action stream). July… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-stereo-rgbd.citrus-fruit-rgbdThis dataset contains
train: 1500 images test: 500 images val: 200 images. Each RGB image also has its corresponding depth file (.npy), and the B (Laplacian-based convexity cues), N (surface bumpness), generated from the depth image.
pointworld-rgbd-stage2RGBDrgbdtable_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.ImageNet_val_RGBD_DepthAnything_SBLprocessed-rgbd-objects-14-classes-uw
Dataset Card for "processed-rgbd-objects-14-classes-uw"
More Information needed
sort_bolt_nut_rgbd_40epsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 40,
"total_frames": 10000,
"total_tasks": 4,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zeuzei/sort_bolt_nut_rgbd_40eps.RGBDT500-mirrorpick-up-erasers-place-in-drawers-rgbd-v2-fixed-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/namin72/pick-up-erasers-place-in-drawers-rgbd-v2-fixed-merged.ImageNet_RGBD_DepthAnything
ImageNet with Depth generated by DepthAnything ViT-L
pick-up-erasers-place-in-drawers-rgbd-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/namin72/pick-up-erasers-place-in-drawers-rgbd-merged.SmolVLA_RGBD_Merged_FORMAL_200_EPISODES_20260805Syn4D_RGBDpiperx-old-ab-rgbd-5903
PiperX Old AB RGB-D — 5,903 accepted points
This public dataset export contains the exact 5,903 accepted training Episodes recorded by the old AB SE(3) perturbation collector. The export is bound to the aggregate state.json truth and contains 5,903 unique schedule_index values across 124 closed shards (shard-00000 through shard-00123). Audit rejections remain audit records and are not training Episodes.
Data
Schema: piperx_lerobot_se3_rgbd_v2
Accepted Episodes /… See the full description on the dataset page: https://huggingface.co/datasets/Travor278/piperx-old-ab-rgbd-5903.quadloco-vla-object_relative-rgbd-1000game-scenes-posed-rgbd
Origin Lab Game Scenes: Posed RGB-D Flythroughs of Game Worlds
Every frame carries the camera that rendered it and the depth the engine computed for it. Ten game worlds, with the camera released from the player for 60% of the footage: metric depth, world-space normals, 4x4 pose, and per-frame intrinsics on one frame index, plus hundreds of full in-place turns and long stretches in which the world is frozen and only the camera moves. Two trajectories per world and one whole… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-scenes-posed-rgbd.COME15K
Dataset Card for "COME15K"
More Information needed
warehouse-rgbd-smolRGPT
SmolRGPT Dataset: Efficient Spatial Reasoning for Warehouse Environments
This repository hosts the Spacial Warehouse Dataset, a key component for the research presented in:
Paper: SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters
Code: https://github.com/abtraore/SmolRGPT
Abstract
Recent advances in vision-language models (VLMs) have enabled powerful multimodal reasoning, but state-of-the-art approaches typically rely on extremely… See the full description on the dataset page: https://huggingface.co/datasets/Abdrah/warehouse-rgbd-smolRGPT.SPAR-Bench-RGBD
🎯 SPAR-Bench-RGBD
A depth-enhanced version of SPAR-Bench for evaluating 3D-aware spatial reasoning in vision-language models.
SPAR-Bench-RGBD extends the full SPAR-Bench with additional depths, camera intrinsics, and pose information, enabling evaluation of models with geometric or 3D-awareness capabilities.The benchmark contains 7,207 manually verified QA pairs across 20 spatial tasks and supports single-view and multi-view inputs.… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhango/SPAR-Bench-RGBD.pick-up-erasers-place-in-drawers-rgbd-v2-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/namin72/pick-up-erasers-place-in-drawers-rgbd-v2-merged.occany-sun-rgbd
SUN RGB-D transfer archive
This repository contains the SUN RGB-D directory used by OccAny. The archive is
split into 512 MiB parts to avoid transferring hundreds of thousands of small
files individually and to make interrupted transfers easy to resume.
Restore
cd /path/to/downloaded/repo
sha256sum -c parts/SHA256SUMS
cat parts/sun-rgbd.tar.zst.part-* > sun-rgbd.tar.zst
sha256sum -c sun-rgbd.tar.zst.sha256
mkdir -p /path/to/SUN_RGB-D
tar --zstd -xf… See the full description on the dataset page: https://huggingface.co/datasets/lr040122/occany-sun-rgbd.
