Panorama-grounding/PanoCaps
PanoCaps: A Human-Annotated Benchmark for Panoptic Grounded Captioning PanoCaps is a benchmark for panoptic grounded captioning: a model writes a full-scene caption and grounds every mentioned entity, things and stuff alike, to pixel-level masks. It contains 3,470 images and 34K panoptic regions, averaging ~9 grounded entities per image, with >99% of regions grounded. Captions are human-written and verified, cover the entire visible scene, use open-vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/Panorama-grounding/PanoCaps.
This repository belongs to Panorama-grounding on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
