datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hot3d
HOT3D-Clips
This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset.
Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here.
See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.).
More details can be found in the HOT3D paper and BOP 2024 report.
hot3d-fullhot3d
HOT3D dataset
Resources: Homepage | HOT3D Toolkit | Whitepaper | Download
Introduction
HOT3D is a dataset for benchmarking egocentric tracking of hands and objects in 3D. The dataset includes 833 minutes of multi-view image streams, which show 19 subjects interacting with 33 diverse rigid objects and are annotated with accurate 3D poses and shapes of hands and objects. HOT3D is recorded with two head-mounted devices from Meta: Project Aria, a research prototype of… See the full description on the dataset page: https://huggingface.co/datasets/projectaria/hot3d.hot3d-annotations-v1
hot3d-annotations-v1
Annotations only — no images, no video.
VITRA-style hand episodes for HOT3D, with per-hand instructions and paraphrases.
episodes
18,805
training samples (index_frame_pair rows)
619,680
annotation
MANO pose + world/camera joints + per-frame extrinsics
text
one instruction per episode + 1.96 paraphrases on average
source frame rate
30 fps
recordings
126 (Aria only)
images / video
not included — see Getting the frames below… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/hot3d-annotations-v1.hot3d
HOT3D EgoHOI clips for the RGB object-entity (OEE) model
Exactly the data that the single-branch + object-entity version of the EgoHOI
Wan2.2-TI2V-5B fine-tune reads at training and evaluation time. Nothing else
from the original working tree is included (no pose caches, depth, ViPE
trajectories or routing boxes); the pseudotactile/ and pointmap/ add-ons
serve the later versions (see their sections below).
Derived from Meta's HOT3D dataset (Aria clips). The HOT3D license and… See the full description on the dataset page: https://huggingface.co/datasets/zhenyuxie-zhzh/hot3d.hot3d_clip_gthot3d-vitra-atomic-instructions-v1
HOT3D VITRA Atomic Instructions v1
This is the language-conditioned, atomic-episode derivative of
MIT-Media-Lab/hot3d-vitra-strict-v1. It contains
32,850 hand-specific episodes, references
250 unique strict source videos, and
indexes 1,227,851 target-hand frames.
Construction
Source: immutable strict-v1 commit
191a139da6dec2e55f73c1558679fd791b19e473.
Left and right hands are segmented independently inside continuous aligned
MANO runs.
Recording-world wrist… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/hot3d-vitra-atomic-instructions-v1.hot3d-videos-formalhot3d_v1hot3d_ithot3d
HOT3D dataset
Resources: Homepage | HOT3D Toolkit | Whitepaper | Download
Introduction
HOT3D is a dataset for benchmarking egocentric tracking of hands and objects in 3D. The dataset includes 833 minutes of multi-view image streams, which show 19 subjects interacting with 33 diverse rigid objects and are annotated with accurate 3D poses and shapes of hands and objects. HOT3D is recorded with two head-mounted devices from Meta: Project Aria, a research prototype… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz043/hot3d.hot3d-videoshot3d_minihot3d-full-zippedHot3D_newcanonical-hot3dmissingfiles_hot3d_3dgs_aria_dynamic
