datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
egocentric-stereo-rgbd
Hub Egocentric: Stereo RGB-D
9 egocentric stereo RGB-D clips with dense metric depth, IMU, 6-DoF VIO pose, and calibration. Each clip carries a self-contained LeRobot v3.0 dataset (loads on lerobot >= 0.6.0) plus side-by-side stereo, mono, a colorized depth preview, and a Foxglove MCAP recording.
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on StereoLabs ZED X Mini. Egocentric, human-demonstration data (passive; no robot action stream). July… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-stereo-rgbd.citrus-fruit-rgbdThis dataset contains
train: 1500 images test: 500 images val: 200 images. Each RGB image also has its corresponding depth file (.npy), and the B (Laplacian-based convexity cues), N (surface bumpness), generated from the depth image.
RGBDtable_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.game-scenes-posed-rgbd
Origin Lab Game Scenes: Posed RGB-D Flythroughs of Game Worlds
Every frame carries the camera that rendered it and the depth the engine computed for it. Ten game worlds, with the camera released from the player for 60% of the footage: metric depth, world-space normals, 4x4 pose, and per-frame intrinsics on one frame index, plus hundreds of full in-place turns and long stretches in which the world is frozen and only the camera moves. Two trajectories per world and one whole… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-scenes-posed-rgbd.piperx-old-ab-rgbd-5903
PiperX Old AB RGB-D — 5,903 accepted points
This public dataset export contains the exact 5,903 accepted training Episodes recorded by the old AB SE(3) perturbation collector. The export is bound to the aggregate state.json truth and contains 5,903 unique schedule_index values across 124 closed shards (shard-00000 through shard-00123). Audit rejections remain audit records and are not training Episodes.
Data
Schema: piperx_lerobot_se3_rgbd_v2
Accepted Episodes /… See the full description on the dataset page: https://huggingface.co/datasets/Travor278/piperx-old-ab-rgbd-5903.COME15K
Dataset Card for "COME15K"
More Information needed
warehouse-rgbd-smolRGPT
SmolRGPT Dataset: Efficient Spatial Reasoning for Warehouse Environments
This repository hosts the Spacial Warehouse Dataset, a key component for the research presented in:
Paper: SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters
Code: https://github.com/abtraore/SmolRGPT
Abstract
Recent advances in vision-language models (VLMs) have enabled powerful multimodal reasoning, but state-of-the-art approaches typically rely on extremely… See the full description on the dataset page: https://huggingface.co/datasets/Abdrah/warehouse-rgbd-smolRGPT.SPAR-Bench-RGBD
🎯 SPAR-Bench-RGBD
A depth-enhanced version of SPAR-Bench for evaluating 3D-aware spatial reasoning in vision-language models.
SPAR-Bench-RGBD extends the full SPAR-Bench with additional depths, camera intrinsics, and pose information, enabling evaluation of models with geometric or 3D-awareness capabilities.The benchmark contains 7,207 manually verified QA pairs across 20 spatial tasks and supports single-view and multi-view inputs.… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhango/SPAR-Bench-RGBD.SPAR-Bench-Tiny-RGBD
🎯 SPAR-Bench-Tiny-RGBD
A lightweight RGBD version of SPAR-Bench for fast evaluation of 3D-aware spatial reasoning in vision-language models (VLMs).
SPAR-Bench-Tiny-RGBD is a subset of SPAR-Bench-RGBD, containing 1,000 QA samples (50 per task × 20 tasks), each augmented with depths, camera intrinsics, and pose information.This dataset is ideal for quick evaluation of 3D-aware models, while maintaining compatibility with the same structure as… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhango/SPAR-Bench-Tiny-RGBD.RoboFactory-5Task-RGBD-Decentralized
RoboFactory Five-Task Decentralized Wrist RGB-D
Public research corpus for reproducible Stereo-CoRE shared-policy experiments.
Contract
500 successful synchronized demonstrations: 100 each of LiftBarrier (2 robots), CameraAlignment (3), ThreeRobotsStackCube (3), LongPipelineDelivery (4), and TakePhoto (4).
A policy stream contains only one panda_hand wrist RGB-D observation and that robot's qpos.
RGB is 640x480; depth is native metric depth stored in millimetres.… See the full description on the dataset page: https://huggingface.co/datasets/B111ue/RoboFactory-5Task-RGBD-Decentralized.RGBD-GSD
RGBD-GSD — RGB-D Glass Surface Detection Dataset
RGBD-GSD is the first large-scale RGB-D glass surface detection dataset, introduced in:
Leveraging RGB-D Data with Cross-Modal Context Mining for Glass Surface DetectionJiaying Lin*, Yuen-Hei Yeung*, Shuquan Ye, Rynson W. H. LauAAAI 2025arXiv · Project Page
Dataset Summary
RGBD-GSD contains 3,009 RGB-D images across a wide range of real-world glass surface categories, each paired with a precise binary segmentation… See the full description on the dataset page: https://huggingface.co/datasets/garrying/RGBD-GSD.RGBD-VideoCount
RGBD-VideoCount
RGBD-VideoCount is an RGB-D video dataset for video object counting in crowded and occluded scenes. It provides synchronized RGB frames and depth maps, together with instance-level annotations for evaluating detection, cross-frame association, and video-level de-duplication.
Dataset Summary
195 RGB-D video clips
6 object categories
2,032 finely annotated frames
77,638 instance bounding boxes
Multi-category shelf and crowded-object scenes
RGB… See the full description on the dataset page: https://huggingface.co/datasets/aerospace123/RGBD-VideoCount.testThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.pointworld-rgbdrgbd_dataset
