CoolFace
Datasetpublic

Aquiles-ai/LLaVA-CC3M-Pretrain-595K-Embedded

Dataset derived from liuhaotian/LLaVA-CC3M-Pretrain-595K Dataset details Dataset type: LLaVA Visual Instruct CC3M Pretrain 595K is a subset of CC-3M dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning. We aim to build large multimodal towards GPT-4 vision/language capability.

sourceHugging Faceupdated 2mo agoView on Hugging Face
2likes247downloads
settings

This repository belongs to Aquiles-ai on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameLLaVA-CC3M-Pretrain-595K-Embedded
visibilitypublic
licencenot set
gatedno
ownerAquiles-ai
Account settings
Aquiles-ai/LLaVA-CC3M-Pretrain-595K-Embedded · CoolFace