datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arkit_labelmaker
ARKit Labelmaker: A New Scale for Indoor 3D Scene Understanding
[arxiv] [website] [checkpoints] [code]
We complement ARKitScenes dataset with dense semantic annotations that are automatically generated at scale. This produces the first large-scale, real-world 3D dataset with dense semantic annotations.
Training on this auto-generated data, we push forward the state-of-the-art performance on ScanNet and ScanNet200 with prevalent 3D semantic segmentation models.
arkitscenes-rrd
ARKitScenes → Rerun (.rrd)
5,015 ARKitScenes indoor iPhone/iPad captures,
converted into layered Rerun recordings — including the data that
exists only inside the dataset's .mov containers and appears in no published asset:
60 Hz camera poses (ARKit visionTransform, 6× denser than the published 10 Hz trajectory) —
validated against the published trajectory per sequence and used when a rigid fit agrees within
3° / 10 cm (pose_source = mebx_stream_4_vision_transform), otherwise… See the full description on the dataset page: https://huggingface.co/datasets/rerun/arkitscenes-rrd.arkitscenes_mcmc_3dgs
Data Statistics
Scenes
Mean PSNR ↑
Mean SSIM ↑
Mean LPIPS ↓
Mean Depth L1 ↓
Mean #3DGS
Total #3DGS
1,290
31.63 dB
0.907
0.217
0.0051 m
1.149M
1.483B
arkitscenes-rrd
ARKitScenes → Rerun (.rrd)
5,015 ARKitScenes indoor iPhone/iPad captures,
converted into layered Rerun recordings — including the data that
exists only inside the dataset's .mov containers and appears in no published asset:
60 Hz camera poses (ARKit visionTransform, 6× denser than the published 10 Hz trajectory) —
validated against the published trajectory per sequence and used when a rigid fit agrees within
3° / 10 cm (pose_source = mebx_stream_4_vision_transform), otherwise… See the full description on the dataset page: https://huggingface.co/datasets/pablovela5620/arkitscenes-rrd.arkit_lowres_processedARKitScenes_preprocessed
ARKitScenes Dataset for LBM
Preprocessed ARKitScenes Dataset following CUT3R.
Each scene contains the following structure when extracted:
40753679/
├── lowres_depth
| ├── 40753679_6790.148.png
| ├── ...
├── vga_wide
| ├── 40753679_6790.148.jpg
| ├── ...
├── new_scene_metadata.npz
└── scene_metadata.npz
Data Format Details:
vga_wide: RGB images in .jpg format
lowres_depth: Depth data in .png format (numpy arrays)
new_scene_metadata.npz: Camera data processed by… See the full description on the dataset page: https://huggingface.co/datasets/ZhengGuangze/ARKitScenes_preprocessed.audio2face-mediapipe-arkit-teacher
audio2face-mediapipe-arkit-teacher
Left: source video frame (face-cropped). Middle: MediaPipe FaceLandmarker's 478 landmark points. Right: an illustrative subset of mp_bs — the 52-channel ARKit blendshape vector shipped in this dataset — as horizontal bars updating per frame.
14,703 emotional-speech clips, each annotated with a 52-channel ARKit blendshape sequence extracted by MediaPipe FaceLandmarker from the source video (or from audio-driven synthesis where no… See the full description on the dataset page: https://huggingface.co/datasets/myned-ai/audio2face-mediapipe-arkit-teacher.arkitscenes-spatiallm
ARKitScenes-SpatialLM Dataset
ARkitScenes dataset preprocessed in SpatialLM format for oriented object bouding boxes detection with LLMs.
Overview
This dataset is derived from ARKitScenes 5,047 real-world indoor scenes captured using Apple's ARKit framework, preprocessed and formatted specifically for SpatialLM training.
Data Extraction
Point clouds and layouts are compressed in zip files. To extract the files, run the following script:
cd arkitscenes-spatiallm… See the full description on the dataset page: https://huggingface.co/datasets/ysmao/arkitscenes-spatiallm.3dscene-arkitsceneARKitScenes_processedThe processed ARKitScenes datasets by VGGT-Det.
Run 'cat train.tar.part_0{00,01,02} > train.tar' to merge the split archives into a single tar file.
Run 'md5sum -c MD5SUMS.txt' to verify that the files were downloaded successfully.
If the dataset is helpful, please cite:
@inproceedings{
dehghan2021arkitscenes,
title={{ARK}itScenes - A Diverse Real-World Dataset for 3D Indoor Scene Understanding Using Mobile {RGB}-D Data},
author={Gilad Baruch and Zhuoyuan Chen and Afshin… See the full description on the dataset page: https://huggingface.co/datasets/YangCaoCS/ARKitScenes_processed.audio2face-emotion-arkit-teacher
audio2face-emotion-arkit-teacher
Nyx avatar (Gaussian-splat head, ARKit-52 blendshape rig) driven by a surprise clip's blendshape labels derived from this dataset.
14,082 emotional-speech clips, each annotated with two parallel 52-channel ARKit blendshape sequences (NVIDIA Audio2Face-3D-v2.3.1-James and LAM_Audio2Expression) plus a 26-dimensional NVIDIA Audio2Emotion conditioning vector.
Reference-only dataset — the original audio is not shipped. Each row contains a… See the full description on the dataset page: https://huggingface.co/datasets/myned-ai/audio2face-emotion-arkit-teacher.arkitscenes-compressedaudio2face-emotion-arkit-teacher
Audio2Face Emotion ARKit Teacher Labels (TsFile)
Apache TsFile version of myned-ai/audio2face-emotion-arkit-teacher.
Overview
14,082 emotional-speech clips, each annotated with two parallel 52-channel ARKit blendshape sequences (NVIDIA Audio2Face-3D-v2.3.1-James and LAM_Audio2Expression) plus a 26-dimensional NVIDIA Audio2Emotion conditioning vector. Released by myned-ai to support teacher-student distillation research: condensing heavy GPU-bound audio2face… See the full description on the dataset page: https://huggingface.co/datasets/THULab/audio2face-emotion-arkit-teacher.arkitscenes-crossview-inpaint
Carlhahaha/arkitscenes-crossview-inpaint
Cross-view pairs generated from inpainting results (Topomap; meta root: /mnt/NAS/data/jz4725/topomap).
Splits
Uploaded splits: test, train, validation
Schema (columns)
Image_a: Original image of sample A (datasets.Image)
Image_b: Original image of sample B (datasets.Image)
Inpaint_b: Inpainted image of sample B (datasets.Image)
mask_b: Mask used for inpainting on B (datasets.Image)
point_b: Normalized centroid of B's… See the full description on the dataset page: https://huggingface.co/datasets/Carlhahaha/arkitscenes-crossview-inpaint.xx_beat_arkit_moshi_2025_07_20_30fps_attn3dscene-arkitscene-highrestis-scenes-arkit-mp3dconcerto_arkitscenes_compressedarkitscenes-preprocessedgauss_gym_arkitsynthetic_accented_englishArkitscenes-Spatiallm
ARKitScenes-SpatialLM Dataset
ARkitScenes dataset preprocessed in SpatialLM format for oriented object bouding boxes detection with LLMs.
Overview
This dataset is derived from ARKitScenes 5,047 real-world indoor scenes captured using Apple's ARKit framework, preprocessed and formatted specifically for SpatialLM training.
Data Extraction
Point clouds and layouts are compressed in zip files. To extract the files, run the following script:
cd arkitscenes-spatiallm… See the full description on the dataset page: https://huggingface.co/datasets/Gen3DF/Arkitscenes-Spatiallm.items_prompts_fullitems_raw_fullxx_beat_arkit_encodec2_2025_07_31_30fps_attnitems_raw_litexx_cbh_arkit_face_v1_moshi_2025_08_10_30fps_attnarkitscenes_processedarkit_scan_procGrupoRTX
