datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hypersimhypersim-frustum-completion
Hypersim Frustum Point Completion
A large-scale indoor point-cloud completion dataset derived from Hypersim. Each record simulates a partially observed room: one real camera view provides context, while a synthetic nearby “missing camera” frustum hides part of the scene. Models are trained to infer the masked region from visible points.
Why this dataset exists. Real 3D capture — whether from depth sensors, multi-view reconstruction, or neural fields — routinely produces… See the full description on the dataset page: https://huggingface.co/datasets/izhleba/hypersim-frustum-completion.Hypersim-ProcessedHypersim-fullhypersimhypersim-examples3dvlm-hypersim
3DVLM Hypersim — §6 depth format
Hypersim (Apple, ICCV 2021) reconverted to the 3DVLM project's unified §6
on-disk format. 457 scenes, ~77,400 frames, V-Ray ground truth.
License: CC BY-SA 3.0, same as the source. If you use this, please cite
Roberts et al. (Hypersim) and apply share-alike terms to any derivatives.
Layout
One uncompressed .tar per scene at the repo root:
hypersim/
ai_001_001.tar
ai_001_002.tar
...
ai_055_010.tar # 457 tars, ~620 MB… See the full description on the dataset page: https://huggingface.co/datasets/helioom/3dvlm-hypersim.hypersim-episodes-v3-parquet
hypersim-episodes-v3-parquet
Per-frame Parquet dataset for ReCAST tracker training.
Schema
One row per frame, grouped by episode_id. Arrow memory-mapped access
enables reading specific frames without loading entire episodes.
Column
Type
Description
episode_id
int32
Episode identifier
frame_idx
int32
Frame index within episode
jpeg
binary
JPEG-encoded RGB frame
depth
list<float32>
Flat H×W depth map
seg
list<uint16>
Semantic segmentation (empty if… See the full description on the dataset page: https://huggingface.co/datasets/OSResight/hypersim-episodes-v3-parquet.Processed_Hypersimhypersim_relative_depth
Hypersim Relative Depth
This repository contains a repackaged and preprocessed version of the
Hypersim Dataset, prepared for relative
depth training with the Marigold V2 codebase.
The repository contains RGB/depth pairs and split metadata organized into
downloadable tar archives. The preprocessing and archive layout are intended
for use with the Marigold V2 data-loading pipeline.
Source dataset
The data originates from:
Dataset: Hypersim
Original repository:… See the full description on the dataset page: https://huggingface.co/datasets/obukhovai/hypersim_relative_depth.layout_diffusion_hypersimThis repository contains the data for SceneCraft: Layout-Guided 3D Scene Generation.
Project page: https://orangesodahub.github.io/SceneCraft
Code: https://github.com/OrangeSodahub/SceneCraft
HyperSim-Absolute-Camera
HyperSim-Absolute-Camera
Per-frame camera parameter annotations for the HyperSim dataset
(448 scenes; 73,598 valid per-frame annotations),
captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).
Format… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/HyperSim-Absolute-Camera.hypersim-trainhypersim-frustum-completion-v2hypersim_mcmc_3dgs
Data Statistics
Scenes
Mean PSNR ↑
Mean SSIM ↑
Mean LPIPS ↓
Mean Depth L1 ↓
Mean #3DGS
Total #3DGS
455
32.51 dB
0.942
0.078
0.0025 m
2.494M
1.135B
License Notice:This dataset is derived from the Hypersim source dataset and follows CC BY-SA 3.0.
3dvlm-hypersim_subsetHypersim_processedpreprocessed_HypersimHypersim-Distractor
Hypersim Synthetic Distractor Pairs
Clean/distractor image pairs for multi-view feature restoration, generated by compositing selected Objaverse objects and their estimated shadows into Hypersim indoor views. This is a derived dataset, not an official Hypersim or Objaverse release.
Release contents
Packaged pairs: 20,965 (train 20,885, eval 80).
Archives: 165; total 42.22 GB (decimal). Inventory updated: 2026-09-14T20:58:19.985376+09:00.
Collection
Pairs… See the full description on the dataset page: https://huggingface.co/datasets/cyjcyj91/Hypersim-Distractor.hypersimHypersim_resizehypersim-mini
Hypersim Minimal (RGB + Semantic Mapped)
This dataset is a minimal extraction from Hypersim:
RGB: preview (either preview JPGs or HDR color.hdf5 tonemapped to PNG)
Semantic labels: mapped to uint8 PNG using clip40to39
Columns
image — RGB image
mask — segmentation mask (uint8)
scene, camera, frame — identifiers
Splits
train
validation (ratio: 0.1)
Note: You are responsible for complying with the original Hypersim license/terms.
hypersimhypersim_big
Hypersim Minimal (RGB + Semantic Mapped)
This dataset is a minimal extraction from Hypersim:
RGB: preview (either preview JPGs or HDR color.hdf5 tonemapped to PNG)
Semantic labels: mapped to uint8 PNG using clip40to39
Columns
image — RGB image
mask — segmentation mask (uint8)
scene, camera, frame — identifiers
Splits
train
validation (ratio: 0.1)
Note: You are responsible for complying with the original Hypersim license/terms.
hypersim_tarHypersim-SpatialLMspatial-inconsistencies-hypersim-curated-615
Hypersim curated spatial-inconsistency benchmark
Private release accompanying Multimodal Language Models Cannot Spot Spatial Inconsistencies.
This repository contains the final manually curated benchmark: 615 pairs across 216 Hypersim scenes. Each pair directory has A.jpg (spatially inconsistent) and B.jpg (matched original). annotations.json provides the evaluation label and metadata.
Provenance and use
The images derive from Hypersim. This private repository is… See the full description on the dataset page: https://huggingface.co/datasets/reachomk/spatial-inconsistencies-hypersim-curated-615.Hypersim_withpchypersim_testHypersim_600_resize_float32
