zzliang/GRIT
GRIT: Large-Scale Training Corpus of Grounded Image-Text Pairs Dataset Summary We introduce GRIT, a large-scale dataset of Grounded Image-Text pairs, which is created based on image-text pairs from COYO-700M and LAION-2B. We construct a pipeline to extract and link text spans (i.e., noun phrases, and referring expressions) in the caption to their corresponding image regions. More details can be found in the paper. Supported Tasks During the… See the full description on the dataset page: https://huggingface.co/datasets/zzliang/GRIT.
161970
update readme
update readme
update readme
update readme
update readme
update readme
update readme
update
add parquet
initial commit
