CoolFace
Datasetpublic

wovenbytoyota-vai/InstVL

InstVL: A Large-Scale Instance-Aware Vision-Language Dataset This is the official repository for the InstVL dataset, introduced in the paper InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding. InstVL is a large-scale dataset of images and videos designed to bridge the gap between holistic scene understanding and fine-grained, instance-level comprehension. Current vision-language pre-training (VLP) paradigms excel at global scene understanding… See the full description on the dataset page: https://huggingface.co/datasets/wovenbytoyota-vai/InstVL.

sourceHugging Faceupdated 6mo agoView on Hugging Face
5likes287downloads

wovenbytoyota-vai/InstVL · main · files are served by the source, never re-hosted here