zzliang/GRIT
GRIT: Large-Scale Training Corpus of Grounded Image-Text Pairs Dataset Summary We introduce GRIT, a large-scale dataset of Grounded Image-Text pairs, which is created based on image-text pairs from COYO-700M and LAION-2B. We construct a pipeline to extract and link text spans (i.e., noun phrases, and referring expressions) in the caption to their corresponding image regions. More details can be found in the paper. Supported Tasks During the… See the full description on the dataset page: https://huggingface.co/datasets/zzliang/GRIT.
This repository belongs to zzliang on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
