datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPIC-Camera
GPIC-Camera
Per-image camera parameter annotations for the GPIC dataset
(train / test / val; train = 8,000 shards, test = 1,000,000 images, val = 200,000 images), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/GPIC-Camera.CLIP-GPIC-embeddings
CLIP-GPIC-embeddings
Vision Transformer embeddings for split:TEST (1M) of stanford-vision-lab/gpic
Mainly for use with cross-attention read/no-read bridge experiments from my github.
Features: pre-indexed image TAR member byte offset and size -> HTTP Range requests to fetch only selected image bytes.
Meaning: Won't require downloading 1M images for retrieval, will just fetch matches from remote shard.
The MIT license applies to the embedding-bank files, metadata… See the full description on the dataset page: https://huggingface.co/datasets/zer0int/CLIP-GPIC-embeddings.
