wovenbytoyota-vai/InstVL
InstVL: A Large-Scale Instance-Aware Vision-Language Dataset This is the official repository for the InstVL dataset, introduced in the paper InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding. InstVL is a large-scale dataset of images and videos designed to bridge the gap between holistic scene understanding and fine-grained, instance-level comprehension. Current vision-language pre-training (VLP) paradigms excel at global scene understanding… See the full description on the dataset page: https://huggingface.co/datasets/wovenbytoyota-vai/InstVL.
5287
Update README.md
Update method title
Update citation
fix identifier
chore: track *.jsonl via Git LFS
initial commit
