tomhodemon/grounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning Code for generating the annotations can be found here: github.com Dataset Summary This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image. Data instance Each sample instance has the following structure: Field Type Description image_file… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.
2347
