spatialvid
SpatialVIDSpatialVID: A Large-Scale Video Dataset with Spatial Annotations
Jiahao Wang1*
Yufeng Yuan1*
Rujie Zheng1*
Youtian Lin1
Jian Gao1
Lin-Zhuo Chen1
Yajie Bao1
Yi Zhang1
Chang Zeng1
Yanxi Zhou1
Xiaoxiao Long1
Hao Zhu1
Zhaoxiang Zhang2
Xun Cao1
Yao Yao1†
1Nanjing University 2Institute of Automation, Chinese Academy of Science
*Equal Contribution †Corresponding Author
CVPR 2026… See the full description on the dataset page: https://huggingface.co/datasets/SpatialVID/SpatialVID.SpatialVID-HQSpatialVID: A Large-Scale Video Dataset with Spatial Annotations
Jiahao Wang1*
Yufeng Yuan1*
Rujie Zheng1*
Youtian Lin1
Jian Gao1
Lin-Zhuo Chen1
Yajie Bao1
Yi Zhang1
Chang Zeng1
Yanxi Zhou1
Xiaoxiao Long1
Hao Zhu1
Zhaoxiang Zhang2
Xun Cao1
Yao Yao1†
1Nanjing University 2Institute of Automation, Chinese Academy of Science
*Equal Contribution †Corresponding Author
CVPR 2026… See the full description on the dataset page: https://huggingface.co/datasets/FelixYuan/SpatialVID-HQ.SpatialVID_HD_InP_encodedspatialvid-trajectory-cluster-media
SpatialVID Trajectory Cluster Review Media
This repository contains 224 selected SpatialVID video clips and matching poster frames used by the SpatialVID Trajectory Clusters Space.
The clips provide up to three visual review examples for every cluster at K=8, 16, 32, 48, and 64. Coarser K values reuse samples from the K=64 selection where possible. Selection is based on released pose geometry and does not certify pose accuracy or capture hardware.
The source videos come from… See the full description on the dataset page: https://huggingface.co/datasets/Rongjin03/spatialvid-trajectory-cluster-media.GimbalDiffusion-SpatialVID-Extreme
SpatialVID-Extreme
The 138-sample extreme-camera-motion benchmark for GimbalDiffusion: Gravity-Aware Camera Control for Video Generation.
Project page
Paper
SpatialVID-HQ source dataset
Contents
spatialvid_extreme.tar: the complete benchmark in one archive, unpacking directly into the layout below.
test_index.json and test_samples/: the canonical 138 prompts, seeds, absolute/relative camera matrices, intrinsics, and source identifiers.
assets/ground_truth_raw/… See the full description on the dataset page: https://huggingface.co/datasets/lefreud/GimbalDiffusion-SpatialVID-Extreme.SpatialVideoDataset_split
