datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLaVAR
LLaVAR Data: Enhanced Visual Instruction Data with Text-Rich Images
More info at LLaVAR project page, Github repo, and paper.
Training Data
Based on the LAION dataset, we collect 422K pretraining data based on OCR results. For finetuning data, we collect 16K high-quality instruction-following data by interacting with langauge-only GPT-4. Note that we also release a larger and more diverse finetuning dataset below (20K), which contains the 16K we used for the paper. The… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/LLaVAR.LLaVA-VSD-120K
LLaVA-VSD-120K
120K Visual Spatial Description dataset for instruction-tuning Large Language-and-Vision Assistant.
Dataset Details
Visual Spatial Description (VSD) aims to generate texts that describe the spatial relationships between objects within images.
Traditional visual spatial relationship classification (VSRC) methods typically output the spatial relationship between two objects in an image,
often neglecting world knowledge and lacking general language… See the full description on the dataset page: https://huggingface.co/datasets/swordli/LLaVA-VSD-120K.LLaVA-Instruct-21K-COCO-SubSet
subset from https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K
train: 21000
val seen: 3000
val unseen: 2100
test: 6000
