4d
Models
All models matching “4d”Datasets
All datasets matching “4d”4DThinker-Training-Data
4DThinker Training Data
This repository contains the training data for 4DThinker, a framework that enables VLMs to "think with 4D" through dynamic latent mental imagery, built upon SpatialVID and DSR_Suite-Data.
Data Structure
data/
├── dift_data.jsonl # DIFT training data (~38K samples)
├── 4drl_data_filtered.jsonl # 4DRL training data (~37K samples)
└── processed_data/ # Video frames & mask overlays
├── <video_id>/
│ ├── frames/… See the full description on the dataset page: https://huggingface.co/datasets/jankin123/4DThinker-Training-Data.MVOIK-4D
HAT4D: Human-Assisted Training for 4D Dynamic Scene Understanding
MVOIK-4D: Multi-View Object Interaction Knowledge for 4D Physical Reasoning
If you find this dataset useful, please consider citing our paper and following the project page for updates.
💡 Description
MVOIK-4D is a curated release of real-world object-interaction sequences for 4D scene understanding, reconstruction, and evaluation. It contains RGB input frames, multi-view evaluation frames, and memory-mask… See the full description on the dataset page: https://huggingface.co/datasets/Lijiaxin0111/MVOIK-4D.4D-Lung
4D-Lung (segmentation subset)
Longitudinal 4D (respiratory-gated, phase-resolved) fan-beam CT of 20
locally-advanced non-small-cell lung cancer (NSCLC) patients, with expert
manual RTSTRUCT contours, from Data from 4D Lung Imaging of NSCLC Patients
(4D-Lung) on The Cancer Imaging Archive (Hugo et al., VCU).
This is the segmentable subset of the full collection — read carefully.
The full TCIA 4D-Lung collection is 183 GB and contains both 4D fan-beam CT
(4D-FBCT, "4DCT") and 4D… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/4D-Lung.rv-4d54f178b0
Internal video store for RuLips-1k
This repository is the working storage behind
levossadtchi/RuLips-1k:
the actual 224×224 clips collected in August–September 2026, kept for our own
training. It is not a curated dataset: folder names are collection machines,
some clips are duplicates, and only RuLips-1k/manifest.jsonl says which files are
part of the release (store.path / store.member / store.batch per clip).
Layout: lips-*/batchNNNNN/<video>/spanNNN_tK_cut.mp4 and bigN/... —… See the full description on the dataset page: https://huggingface.co/datasets/levossadtchi/rv-4d54f178b0.4DNeX-10M
4DNeX-10M Dataset
📄 Paper | 🚀 Project Page | 💻 GitHub
Introduction
4DNeX-10M is a large-scale hybrid dataset introduced in the paper "4DNeX: Feed-Forward 4D Generative Modeling Made Easy".
The dataset aggregates monocular videos from diverse sources, including both static and dynamic scenes, accompanied by high-quality pseudo 4D annotations generated using state-of-the-art 3D and 4D reconstruction methods. The dataset enables joint modeling of RGB appearance and… See the full description on the dataset page: https://huggingface.co/datasets/3DTopia/4DNeX-10M.iPhone360-4dgs360
iPhone360 Dataset - 4dgs360 preprocessed version
iPhone360 is a benchmark dataset for 360° reconstruction of dynamic objects from monocular video, introduced in the paper:
4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
Jae Won Jang, Yeonjin Chang, Wonsik Shin, Juhwan Cho, Nojun Kwak
Project Page · arXiv
Dataset Description
iPhone360 features real-world dynamic scenes captured with an iPhone, where test cameras are positioned at… See the full description on the dataset page: https://huggingface.co/datasets/mipal/iPhone360-4dgs360.
