KangLiao/DL3DV-Absolute-Camera
DL3DV-Absolute-Camera Per-frame camera parameter annotations for the DL3DV dataset (6,377 scenes across 7 buckets 1K–7K; 2,161,003 valid per-frame annotations), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website. The collage above visualizes the camera maps on sample images — each pair shows the up field (green arrows: the projected gravity-up direction) and the latitude field (colored contours: angle above/below the horizon).… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/DL3DV-Absolute-Camera.
DL3DV-Absolute-Camera
Per-frame camera parameter annotations for the DL3DV dataset (6,377 scenes across 7 buckets 1K–7K; 2,161,003 valid per-frame annotations), captioned by the **Puffin-World** model. More captioned datasets are provided in our **Puffin-16M** website.
The collage above visualizes the camera maps on sample images — each pair shows the up field (green arrows: the projected gravity-up direction) and the latitude field (colored contours: angle above/below the horizon).
Format
The archive mirrors the source DL3DV-ALL-960P layout: one `.zip` per scene, grouped into bucket folders — <bucket>/<scene_hash>.zip (e.g. 1K/001dccbc…740d5f.zip). The scene hashes match DL3DV-ALL-960P, so the camera annotations pair 1:1 with the source frames.
Each zip unpacks to the original per-frame layout:
dense/camera/frame_00001.json
dense/camera/frame_00002.json
...Each JSON holds the predicted monocular camera parameters for that frame:
Example (dense/camera/frame_00001.json):
{"roll": 0.0044, "pitch": 0.1566, "vfov": 1.0864, "k1": 0.0, "parse_ok": true}Camera Parameter Distributions
Histograms of the predicted roll / pitch / vertical-FoV over all frames (proportion of valid samples per 10° bin; parse_ok=False samples excluded).
Reading the distributions
- Roll is tightly peaked at 0° (σ ≈ 4.9°) — DL3DV captures are shot level.
- Pitch carries a downward bias (median ≈ −6°, mean ≈ −7°) with a wide spread (σ ≈ 16°), reflecting the varied up/down camera motion across scenes.
- FoV concentrates around 50–55° (median ≈ 55°) — moderately wide video capture, ranging ~20–90°.
If you'd like a dataset with a more diverse and uniform distribution of camera parameters, please refer to our Puffin-4M and Puffin-16M datasets.
Dataset Download
You can download the entire dataset using the following command:
hf download KangLiao/DL3DV-Absolute-Camera --repo-type datasetFrom Camera Parameters to Up and Latitude Fields
The released (roll, pitch, vfov, k1) annotations can be converted into the dense perspective-field representation used by Puffin-World. The conversion computes focal length from vfov, constructs a radial camera, and maps roll and pitch to the gravity direction. `get_perspective_field` then returns a normalized 2-channel up field and a 1-channel latitude field. The up field encodes the projected world-up direction at every pixel, while the latitude field measures each viewing ray's angular elevation relative to the horizon.
import json
import torch
from scripts.camera.geometry.camera import SimpleRadial
from scripts.camera.geometry.gravity import Gravity
from scripts.camera.geometry.perspective_fields import get_perspective_field
from scripts.camera.utils.conversions import fov2focal
with open("camera.json") as f:
annotation = json.load(f)
roll, pitch, vfov, k1 = (
annotation[key] for key in ("roll", "pitch", "vfov", "k1")
)
H, W = 512, 512
f = float(fov2focal(torch.tensor(vfov), H))
camera = SimpleRadial(torch.tensor(
[W, H, f, f, W / 2, H / 2, k1, 0.0]
).float()).scale(torch.tensor([1.0, 1.0]))
gravity = Gravity.from_rp(torch.tensor(roll), torch.tensor(pitch))
up_field, latitude_field = get_perspective_field(camera, gravity)
# Shapes: [1, 2, H, W] and [1, 1, H, W]Run the example from the Puffin-World directory. Use `plot_vector_fields` and `plot_latitudes` to render up-field arrows and latitude heatmaps or contours. See `save_pf_visualization` for an end-to-end visualization example.
Caption Pipeline
Beyond this captioned dataset, we also release a complete captioning pipeline for annotating camera parameters for arbitrary datasets, analyzing camera parameter distributions, and visualizing the corresponding camera maps. The pipeline is available in our GitHub repository.
Citation
If you find the captioned dataset useful for your research or applications, please cite the following papers using these BibTeX entries:
@article{liao2025puffin,
title={Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation},
author={Liao, Kang and Wu, Size and Wu, Zhonghua and Jin, Linyi and Wang, Chao and Wang, Yikai and Wang, Fei and Li, Wei and Loy, Chen Change},
journal={arXiv preprint arXiv:2510.08673},
year={2025}
}
@article{liao2026puffinworld,
title={Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States},
author={Liao, Kang and Luo, Yihang and Wu, Xiao-Ming and Jin, Linyi and Wu, Size and Lin, Chunyu and Zhao, Yao and Wang, Fei and Li, Wei and Loy, Chen Change},
journal={arXiv preprint arXiv:2609.04196},
year={2026}
}