datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLaVAR
LLaVAR Data: Enhanced Visual Instruction Data with Text-Rich Images
More info at LLaVAR project page, Github repo, and paper.
Training Data
Based on the LAION dataset, we collect 422K pretraining data based on OCR results. For finetuning data, we collect 16K high-quality instruction-following data by interacting with langauge-only GPT-4. Note that we also release a larger and more diverse finetuning dataset below (20K), which contains the 16K we used for the paper. The… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/LLaVAR.K-prism
K-Prism
Korean diagnostic benchmark data for evaluating hallucination in vision-language
models. Evaluation code and protocol documentation are available at
alsgur0720/K-Prism.
Files
File
Contents
Text_track.json
504 text-track questions
Image_track.json
498 image-track questions
images/
165 original images referenced by the image track
The release contains 1,002 questions and approximately 303 MB of annotations and
images. Keep both JSON files… See the full description on the dataset page: https://huggingface.co/datasets/KETI-NLP/K-prism.moma-lrg-state-transitionVL3-Syn7M
The re-caption dataset used in VideoLLaMA 3: Frontier Multimodal Foundation Models for Video Understanding
If you like our project, please give us a star ⭐ on Github for the latest update.
🌟 Introduction
This dataset is the re-captioned data we used during the training of VideoLLaMA3. It consists of 7 million diverse, high-quality images, each accompanied by a short caption and a detailed caption.
The images in this dataset originate from COYO-700M, MS-COCO 2017… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/VL3-Syn7M.ExpArt
Dataset Card for Explain Artworks: ExpArt
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Card for "Wiki-ImageReview1.0"
Dataset Summary
Explain Artworks: ExpArt is designed to enhance the capabilities of large-scale vision-language models (LVLMs) in analyzing and describing artworks.
Drawing from a comprehensive array of English Wikipedia art articles, the dataset encourages LVLMs to… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/ExpArt.PVSG-state-transitionHodHod
