KangLiao/CC12M-Camera
CC12M-Camera Per-image camera parameter annotations for the CC12M (Conceptual 12M) dataset (~10.97M images across 2,176 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website. The collage above visualizes the camera maps on sample images — each pair shows the up field (green arrows: the projected gravity-up direction) and the latitude field (colored contours: angle above/below the horizon). Format One .tar per… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/CC12M-Camera.
CC12M-Camera
Per-image camera parameter annotations for the CC12M (Conceptual 12M) dataset (~10.97M images across 2,176 shards), captioned by the **Puffin-World** model. More captioned datasets are provided in our **Puffin-16M** website.
The collage above visualizes the camera maps on sample images — each pair shows the up field (green arrows: the projected gravity-up direction) and the latitude field (colored contours: angle above/below the horizon).
Format
One .tar per source shard (0000.tar … 2175.tar), each containing one .json per image whose name matches the source image key (shard-local index). Source images: pixparse/cc12m-wds (same 2,176-shard layout / keys).
Each JSON holds the predicted monocular camera parameters:
Example:
{"roll": 0.0123, "pitch": -0.0871, "vfov": 1.0123, "k1": 0.0000, "parse_ok": true}Camera Parameter Distributions
Histograms of the predicted roll / pitch / vertical-FoV over the whole dataset (proportion of valid samples per 10° bin; parse_ok=False excluded).
- Roll is sharply peaked at 0° (web images shot upright/level).
- Pitch is centered close to 0° with a slight downward bias.
- FoV concentrates in 20–45° (median ≈ 29°), with a long wide-angle tail.
If you'd like a dataset with a more diverse and uniform distribution of camera parameters, please refer to our Puffin-4M and Puffin-16M datasets.
Dataset Download
You can download the entire dataset using the following command:
hf download KangLiao/CC12M-Camera --repo-type datasetFrom Camera Parameters to Up and Latitude Fields
The released (roll, pitch, vfov, k1) annotations can be converted into the dense perspective-field representation used by Puffin-World. The conversion computes focal length from vfov, constructs a radial camera, and maps roll and pitch to the gravity direction. `get_perspective_field` then returns a normalized 2-channel up field and a 1-channel latitude field. The up field encodes the projected world-up direction at every pixel, while the latitude field measures each viewing ray's angular elevation relative to the horizon.
import json
import torch
from scripts.camera.geometry.camera import SimpleRadial
from scripts.camera.geometry.gravity import Gravity
from scripts.camera.geometry.perspective_fields import get_perspective_field
from scripts.camera.utils.conversions import fov2focal
with open("camera.json") as f:
annotation = json.load(f)
roll, pitch, vfov, k1 = (
annotation[key] for key in ("roll", "pitch", "vfov", "k1")
)
H, W = 512, 512
f = float(fov2focal(torch.tensor(vfov), H))
camera = SimpleRadial(torch.tensor(
[W, H, f, f, W / 2, H / 2, k1, 0.0]
).float()).scale(torch.tensor([1.0, 1.0]))
gravity = Gravity.from_rp(torch.tensor(roll), torch.tensor(pitch))
up_field, latitude_field = get_perspective_field(camera, gravity)
# Shapes: [1, 2, H, W] and [1, 1, H, W]Run the example from the Puffin-World directory. Use `plot_vector_fields` and `plot_latitudes` to render up-field arrows and latitude heatmaps or contours. See `save_pf_visualization` for an end-to-end visualization example.
Caption Pipeline
Beyond this captioned dataset, we also release a complete captioning pipeline for annotating camera parameters for arbitrary datasets, analyzing camera parameter distributions, and visualizing the corresponding camera maps. The pipeline is available in our GitHub repository.
Citation
If you find the captioned dataset useful for your research or applications, please cite the following papers using these BibTeX entries:
@article{liao2025puffin,
title={Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation},
author={Liao, Kang and Wu, Size and Wu, Zhonghua and Jin, Linyi and Wang, Chao and Wang, Yikai and Wang, Fei and Li, Wei and Loy, Chen Change},
journal={arXiv preprint arXiv:2510.08673},
year={2025}
}
@article{liao2026puffinworld,
title={Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States},
author={Liao, Kang and Luo, Yihang and Wu, Xiao-Ming and Jin, Linyi and Wu, Size and Lin, Chunyu and Zhao, Yao and Wang, Fei and Li, Wei and Loy, Chen Change},
journal={arXiv preprint arXiv:2609.04196},
year={2026}
}