CoolFace
Datasetpublic

KangLiao/CC12M-Camera

CC12M-Camera Per-image camera parameter annotations for the CC12M (Conceptual 12M) dataset (~10.97M images across 2,176 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website. The collage above visualizes the camera maps on sample images — each pair shows the up field (green arrows: the projected gravity-up direction) and the latitude field (colored contours: angle above/below the horizon). Format One .tar per… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/CC12M-Camera.

sourceHugging Faceupdated 16d agoView on Hugging Face
1likes668downloads
Dataset Card

CC12M-Camera

[image]

Per-image camera parameter annotations for the CC12M (Conceptual 12M) dataset (~10.97M images across 2,176 shards), captioned by the **Puffin-World** model. More captioned datasets are provided in our **Puffin-16M** website.

The collage above visualizes the camera maps on sample images — each pair shows the up field (green arrows: the projected gravity-up direction) and the latitude field (colored contours: angle above/below the horizon).

Format

One .tar per source shard (0000.tar2175.tar), each containing one .json per image whose name matches the source image key (shard-local index). Source images: pixparse/cc12m-wds (same 2,176-shard layout / keys).

Each JSON holds the predicted monocular camera parameters:

FieldMeaningUnit
rollcamera rollradians
pitchcamera pitchradians
vfovvertical field-of-viewradians
k1radial distortion coefficient
parse_okwhether the model output parsed within valid rangesbool

Example:

json
{"roll": 0.0123, "pitch": -0.0871, "vfov": 1.0123, "k1": 0.0000, "parse_ok": true}

Camera Parameter Distributions

Histograms of the predicted roll / pitch / vertical-FoV over the whole dataset (proportion of valid samples per 10° bin; parse_ok=False excluded).

[image]

splitroll μ / med / σpitch μ / med / σFoV μ / med / σ
all (10.90M)0.2° / 0.0° / 5.7°−2.8° / −0.3° / 11.8°32.3° / 29.0° / 10.6°
  • Roll is sharply peaked at 0° (web images shot upright/level).
  • Pitch is centered close to 0° with a slight downward bias.
  • FoV concentrates in 20–45° (median ≈ 29°), with a long wide-angle tail.

If you'd like a dataset with a more diverse and uniform distribution of camera parameters, please refer to our Puffin-4M and Puffin-16M datasets.

Dataset Download

You can download the entire dataset using the following command:

bash
hf download KangLiao/CC12M-Camera --repo-type dataset

From Camera Parameters to Up and Latitude Fields

The released (roll, pitch, vfov, k1) annotations can be converted into the dense perspective-field representation used by Puffin-World. The conversion computes focal length from vfov, constructs a radial camera, and maps roll and pitch to the gravity direction. `get_perspective_field` then returns a normalized 2-channel up field and a 1-channel latitude field. The up field encodes the projected world-up direction at every pixel, while the latitude field measures each viewing ray's angular elevation relative to the horizon.

python
import json
import torch
from scripts.camera.geometry.camera import SimpleRadial
from scripts.camera.geometry.gravity import Gravity
from scripts.camera.geometry.perspective_fields import get_perspective_field
from scripts.camera.utils.conversions import fov2focal

with open("camera.json") as f:
    annotation = json.load(f)
roll, pitch, vfov, k1 = (
    annotation[key] for key in ("roll", "pitch", "vfov", "k1")
)

H, W = 512, 512
f = float(fov2focal(torch.tensor(vfov), H))
camera = SimpleRadial(torch.tensor(
    [W, H, f, f, W / 2, H / 2, k1, 0.0]
).float()).scale(torch.tensor([1.0, 1.0]))
gravity = Gravity.from_rp(torch.tensor(roll), torch.tensor(pitch))
up_field, latitude_field = get_perspective_field(camera, gravity)
# Shapes: [1, 2, H, W] and [1, 1, H, W]

Run the example from the Puffin-World directory. Use `plot_vector_fields` and `plot_latitudes` to render up-field arrows and latitude heatmaps or contours. See `save_pf_visualization` for an end-to-end visualization example.

Caption Pipeline

Beyond this captioned dataset, we also release a complete captioning pipeline for annotating camera parameters for arbitrary datasets, analyzing camera parameter distributions, and visualizing the corresponding camera maps. The pipeline is available in our GitHub repository.

Citation

If you find the captioned dataset useful for your research or applications, please cite the following papers using these BibTeX entries:

bibtex
  @article{liao2025puffin,
    title={Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation},
    author={Liao, Kang and Wu, Size and Wu, Zhonghua and Jin, Linyi and Wang, Chao and Wang, Yikai and Wang, Fei and Li, Wei and Loy, Chen Change},
    journal={arXiv preprint arXiv:2510.08673},
    year={2025}
  }

  @article{liao2026puffinworld,
    title={Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States},
    author={Liao, Kang and Luo, Yihang and Wu, Xiao-Ming and Jin, Linyi and Wu, Size and Lin, Chunyu and Zhao, Yao and Wang, Fei and Li, Wei and Loy, Chen Change},
    journal={arXiv preprint arXiv:2609.04196},
    year={2026}
  }