datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VG150-coco-format
VG150 — Visual Genome 150 (COCO format)
This dataset is the standard VG150 split of
Visual Genome
(Krishna et al., 2017), the most widely used benchmark for Scene Graph Generation,
reformatted in standard COCO-JSON format. VG150 contains the top 150 object categories
and 50 relations from the original Visual Genome dataset, selected by frequency in the
Scene Graph Generation by Iterative Message Passing paper.
This version in COCO format was produced as part of the… See the full description on the dataset page: https://huggingface.co/datasets/maelic/VG150-coco-format.coco_captioning_complete_formatGQA200-coco-format
GQA — General Question Answering (COCO format)
This dataset is the GQA200 split of
the GQA dataset
(Hudson et al., 2019), reformatted in standard COCO-JSON format.
GQA200 contains the top 200 object categories
and 100 relations from the original GQA dataset, selected by frequency in the
Stacked hybrid-attention and group collaborative learning for unbiased scene graph generation
paper. This dataset has no official test split since it was used
for question answering rather than… See the full description on the dataset page: https://huggingface.co/datasets/maelic/GQA200-coco-format.coco-captions-2017-lance
COCO Captions 2017 (Lance Format)
A Lance-formatted version of the COCO Captions 2017 corpus, redistributed via lmms-lab/COCO-Caption2017. Each row is one image with 5–7 human-written captions, a cosine-normalized CLIP image embedding, and a cosine-normalized CLIP text embedding of the canonical caption — all stored inline and available directly from the Hub at hf://datasets/lance-format/coco-captions-2017-lance/data.
Key features
Inline JPEG bytes in the image… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/coco-captions-2017-lance.coco-detection-2017-lance
COCO 2017 Object Detection (Lance Format)
A Lance-formatted version of the COCO 2017 object detection benchmark, sourced from detection-datasets/coco. Each row is one image with its inline JPEG bytes, the full per-image list of bounding boxes, COCO 80-class category ids and names, per-object areas, an OpenCLIP image embedding, and pre-built indices — all available directly from the Hub at hf://datasets/lance-format/coco-detection-2017-lance/data.
Key features
Inline… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/coco-detection-2017-lance.PSG-coco-format
PSG — Panoptic Scene Graph (COCO format)
This dataset is a reformatted version of the Panoptic Scene Graph (PSG) benchmark
(Yang et al., NeurIPS 2022) in standard COCO-JSON
format, ready for use with object detection and scene graph generation pipelines.
It was produced as part of the
SGG-Benchmark framework and used to train
the models described in the REACT paper
(Neau et al., BMVC 2025).
/!\ Disclaimer: this dataset does NOT contain original segmentation masks, but only
bounding… See the full description on the dataset page: https://huggingface.co/datasets/maelic/PSG-coco-format.IndoorVG-coco-format
IndoorVG — Indoor Visual Genome (COCO format)
IndoorVG is a curated split of
Visual Genome
targeting real-world indoor scenarios (kitchens, offices, living rooms, …).
It was proposed in
Neau et al. (2024)
and reformatted here in standard COCO-JSON format.
It was produced as part of the
SGG-Benchmark framework and used to train
the models described in the REACT paper
(Neau et al., BMVC 2025).
The 84 object classes and 37 predicate classes were manually selected and
semi-automatically… See the full description on the dataset page: https://huggingface.co/datasets/maelic/IndoorVG-coco-format.crater-boulder-moon-coco-format-smallcrater-boulder-moon-coco-format
