CoolFace
Datasetpublic

KangLiao/DL3DV-Absolute-Camera

DL3DV-Absolute-Camera Per-frame camera parameter annotations for the DL3DV dataset (6,377 scenes across 7 buckets 1K–7K; 2,161,003 valid per-frame annotations), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website. The collage above visualizes the camera maps on sample images — each pair shows the up field (green arrows: the projected gravity-up direction) and the latitude field (colored contours: angle above/below the horizon).… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/DL3DV-Absolute-Camera.

sourceHugging Faceupdated 16d agoView on Hugging Face
0likes358downloads
Dataset Card

DL3DV-Absolute-Camera

[image]

Per-frame camera parameter annotations for the DL3DV dataset (6,377 scenes across 7 buckets 1K7K; 2,161,003 valid per-frame annotations), captioned by the **Puffin-World** model. More captioned datasets are provided in our **Puffin-16M** website.

The collage above visualizes the camera maps on sample images — each pair shows the up field (green arrows: the projected gravity-up direction) and the latitude field (colored contours: angle above/below the horizon).

Format

The archive mirrors the source DL3DV-ALL-960P layout: one `.zip` per scene, grouped into bucket folders — <bucket>/<scene_hash>.zip (e.g. 1K/001dccbc…740d5f.zip). The scene hashes match DL3DV-ALL-960P, so the camera annotations pair 1:1 with the source frames.

Each zip unpacks to the original per-frame layout:

dense/camera/frame_00001.json
dense/camera/frame_00002.json
...

Each JSON holds the predicted monocular camera parameters for that frame:

fieldmeaningunit
rollcamera rollradians
pitchcamera pitchradians
vfovvertical field of viewradians
k1radial distortion coefficient
parse_okwhether the model output parsed within valid rangesbool

Example (dense/camera/frame_00001.json):

json
{"roll": 0.0044, "pitch": 0.1566, "vfov": 1.0864, "k1": 0.0, "parse_ok": true}

Camera Parameter Distributions

Histograms of the predicted roll / pitch / vertical-FoV over all frames (proportion of valid samples per 10° bin; parse_ok=False samples excluded).

[image]

splitroll μ / med / σpitch μ / med / σFoV μ / med / σ
all (2.16M frames)−0.3° / −0.3° / 4.9°−7.0° / −6.1° / 16.0°53.4° / 54.8° / 8.7°

Reading the distributions

  • Roll is tightly peaked at 0° (σ ≈ 4.9°) — DL3DV captures are shot level.
  • Pitch carries a downward bias (median ≈ −6°, mean ≈ −7°) with a wide spread (σ ≈ 16°), reflecting the varied up/down camera motion across scenes.
  • FoV concentrates around 50–55° (median ≈ 55°) — moderately wide video capture, ranging ~20–90°.

If you'd like a dataset with a more diverse and uniform distribution of camera parameters, please refer to our Puffin-4M and Puffin-16M datasets.

Dataset Download

You can download the entire dataset using the following command:

bash
hf download KangLiao/DL3DV-Absolute-Camera --repo-type dataset

From Camera Parameters to Up and Latitude Fields

The released (roll, pitch, vfov, k1) annotations can be converted into the dense perspective-field representation used by Puffin-World. The conversion computes focal length from vfov, constructs a radial camera, and maps roll and pitch to the gravity direction. `get_perspective_field` then returns a normalized 2-channel up field and a 1-channel latitude field. The up field encodes the projected world-up direction at every pixel, while the latitude field measures each viewing ray's angular elevation relative to the horizon.

python
import json
import torch
from scripts.camera.geometry.camera import SimpleRadial
from scripts.camera.geometry.gravity import Gravity
from scripts.camera.geometry.perspective_fields import get_perspective_field
from scripts.camera.utils.conversions import fov2focal

with open("camera.json") as f:
    annotation = json.load(f)
roll, pitch, vfov, k1 = (
    annotation[key] for key in ("roll", "pitch", "vfov", "k1")
)

H, W = 512, 512
f = float(fov2focal(torch.tensor(vfov), H))
camera = SimpleRadial(torch.tensor(
    [W, H, f, f, W / 2, H / 2, k1, 0.0]
).float()).scale(torch.tensor([1.0, 1.0]))
gravity = Gravity.from_rp(torch.tensor(roll), torch.tensor(pitch))
up_field, latitude_field = get_perspective_field(camera, gravity)
# Shapes: [1, 2, H, W] and [1, 1, H, W]

Run the example from the Puffin-World directory. Use `plot_vector_fields` and `plot_latitudes` to render up-field arrows and latitude heatmaps or contours. See `save_pf_visualization` for an end-to-end visualization example.

Caption Pipeline

Beyond this captioned dataset, we also release a complete captioning pipeline for annotating camera parameters for arbitrary datasets, analyzing camera parameter distributions, and visualizing the corresponding camera maps. The pipeline is available in our GitHub repository.

Citation

If you find the captioned dataset useful for your research or applications, please cite the following papers using these BibTeX entries:

bibtex
  @article{liao2025puffin,
    title={Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation},
    author={Liao, Kang and Wu, Size and Wu, Zhonghua and Jin, Linyi and Wang, Chao and Wang, Yikai and Wang, Fei and Li, Wei and Loy, Chen Change},
    journal={arXiv preprint arXiv:2510.08673},
    year={2025}
  }

  @article{liao2026puffinworld,
    title={Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States},
    author={Liao, Kang and Luo, Yihang and Wu, Xiao-Ming and Jin, Linyi and Wu, Size and Lin, Chunyu and Zhao, Yao and Wang, Fei and Li, Wei and Loy, Chen Change},
    journal={arXiv preprint arXiv:2609.04196},
    year={2026}
  }