datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLaVA-OneVision-2-Data
LLaVA-OneVision-2-Data
Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training.
At a Glance
The dataset is split across two Hugging Face repositories because of its size:
Repository
What it contains
Part 1 (this repository)
~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.LLaVA-OneVision-1.5-Instruct-Data
LLaVA-OneVision-1.5 Instruction Data
Paper | Code
📌 Introduction
This dataset, LLaVA-OneVision-1.5-Instruct, was collected and integrated during the development of LLaVA-OneVision-1.5. LLaVA-OneVision-1.5 is a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. This meticulously curated 22M instruction dataset (LLaVA-OneVision-1.5-Instruct) is part of a… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Instruct-Data.LLaVA-Video-178K
Dataset Card for LLaVA-Video-178K
Uses
This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy.
Data Sources
For the training of LLaVA-Video, we utilized video-language data from five primary sources:
LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K.LLaVA-OneVision-Data
Dataset Card for LLaVA-OneVision
[2024-09-01]: Uploaded VisualWebInstruct(filtered), it's used in OneVision Stage
almost all subsets are uploaded with HF's required format and you can use the recommended interface to download them and follow our code below to convert them.
the subset of ureader_kg and ureader_qa are uploaded with the processed jsons and tar.gz of image folders.
You may directly download them from the following url.… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Data.LLaVA-ReCap-CC12MLLaVA-OneVision-1.5-Instruct-Data-qwen-formatllava-bench-in-the-wild
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of LLaVA-Bench(wild) that is used in LLaVA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@misc{liu2023improvedllava,
author={Liu, Haotian and Li, Chunyuan… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/llava-bench-in-the-wild.LLaVA-OneVision-Data-ru
LLaVA-OneVision-Data-ru
Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate.
Almost all datasets have been translated, except for the following:
["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"]
Usage
import datasets
data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.LLaVA-NeXT-Data
Dataset Card for LLaVA-NeXT
We provide the whole details of LLaVA-NeXT Dataset. In this dataset, we include the data that was used in the instruction tuning stage for LLaVA-NeXT and LLaVA-NeXT(stronger).
Aug 30, 2024: We update the dataset with raw format (de-compress it for json file and images with structured folder), you can directly download them if you are familiar with LLaVA data format.
Dataset Sources
Compared to the instruction data mixture for LLaVA-1.5… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-NeXT-Data.llava_v1_5_mix665k
LLaVA v1.5 Mix 665K Dataset
This dataset contains 665,298 multimodal instruction-following samples used for fine-tuning the LLaVA v1.5 model.
Dataset Structure
id: Unique identifier for the sample
model: Model name (if applicable)
conversations: JSON string containing conversation turns in original format
image: List of PIL Image objects (embedded in parquet)
image_path: List of strings containing original relative paths to images
Load the Dataset
from… See the full description on the dataset page: https://huggingface.co/datasets/Icey444/llava_v1_5_mix665k.LLaVA-NeXT-Interleave-Bench
LLaVA-Interleave Bench Dataset Card
Dataset details
Dataset type:
LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API.
It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs.
Dataset date:
LLaVA-Interleave Bench was collected in April 2024, and released in June 2024.
Paper or resources for more information:
Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.llava-instruct-mix
LLaVA Instruct Mix
Added OCR and Chart QA dataset into this for more text extraction questions
Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset
Agri-LLaVA
Agri-LLaVA is a large multimodal instruction dataset for agriculture, pairing crop/leaf images with multi-turn diagnostic conversations about plant diseases, pests, and nutrient deficiencies. It is compiled from 16 public source datasets (see the license table below).
This dataset has been converted to Parquet format with image bytes embedded directly, standardized to the HF image_text_to_text format with a single conversational messages schema.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset.LLaVA-ReCap-CC3MLLaVA-CoT-100k
Dataset Card for LLaVA-CoT
The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k.llava-finetunellava-instruct-mix-vsfttheblackcat102/llava-instruct-mix reformated for VSFT with TRL's SFT Trainer.
See https://github.com/huggingface/trl/blob/main/examples/scripts/vsft_llava.py.
llava-665k
LLaVA-v1.5 Mix665K — Arrow (images embedded)
The LLaVA-v1.5 visual instruction-tuning mixture (llava_v1_5_mix665k) converted to a 🤗 datasets
Arrow dataset with image bytes embedded. 665,298 examples
(624,610 image–text + 40,688 text-only).
⚠️ This repo is a raw save_to_disk Arrow snapshot. The dataset viewer and
load_dataset() do not work here — load it with load_from_disk as shown below.
Loading
from huggingface_hub import snapshot_download
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Ethlake/llava-665k.coda-lm-llava-format
CODA-LM Dataset Card
CODA-LM is the multi-modal version of the CODA dataset, used in the CODA-LM paper. Both English and Chinese annotations are available. Check detailed usage in our Github repo.
This repo contains the CODA-LM dataset, which has been reorganized in the LLaVA data format.
You are also welcome to check the original CODA-LM data which contains more metadata vanilla annotations.
Usage
from datasets import load_dataset
# name can be selected from… See the full description on the dataset page: https://huggingface.co/datasets/KaiChen1998/coda-lm-llava-format.llava-en-zh-300kThis dataset is composed by
150k examples of English Visual Instruction Data from LLaVA.
150k examples of English Visual Instruction Data from openbmb.
You can use it in LLaMA Factory by specifying --dataset llava_150k_en,llava_150k_zh.
llava_trajectoriesmidjourney-v6-llavaThis dataset based on https://huggingface.co/datasets/CortexLM/midjourney-v6 dataset, captioned with LLava-1.6 model.
This dataset was released as is. By accessing and using this dataset, you acknowledge and agree that Cortex Foundation and the author of this repo are not responsible for any copyright violations or legal consequences that may arise from the use of these images.
LLaVA-Instruct-600K-Chinese
仿照 LLaVA-Instruct-150K ,使用 Qwen2.5-VL-32B-Instruct 合成的用于微调中文VLM的数据;也可以与英文数据集混合使用,训练多语言VLM
任务类型为基于单张图片的问答和对话,每个样本都对应一张不同的图片,其中大部分图片包含中文字符,更适合中文场景下视觉语言模型的训练。
图片从各类中文网站上爬取
包含3类任务:日常对话、复杂推理、描述图片。日常对话通常是5轮对话,其余任务是1轮对话。
每种任务的数量如下:
任务类型
数量
日常对话
247,431
复杂推理
194,646
描述图片
199,791
用于生成对话数据的prompt如下
日常对话
设计一个你和一个询问这张照片的人之间的对话。答案应该是视觉AI助手看到图像并回答问题的语气。
你需要提出不同的问题并给出相应的答案。问题可以包括询问图像视觉内容的问题,包括对象类型、对象计数、对象动作、对象位置、对象之间的相对位置等。必须是有明确答案的问题,即
(1) 人们可以在图像中明确看到问题所问的内容,并且可以自信地回答;
(2)… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/LLaVA-Instruct-600K-Chinese.llava-instruct-mix
LLaVA Instruct Mix
Summary
The LLaVA Instruct Mix dataset is a processed version of LLaVA Instruct Mix.
Data Structure
Format: Conversational
Type: Language-modeling
Columns:
"images": The image associated with the text.
"prompt": A list of messages that form the context for the conversation.
"completion": The last message in the conversation, which is the model's response.
This structure allows models to learn from the context of the conversation… See the full description on the dataset page: https://huggingface.co/datasets/trl-lib/llava-instruct-mix.LLaVA-Med-60K-IM-text
LLaVA-Med-60K-IM-text
This dataset is a text format of llava_med_instruct_60k_inline_mention.json.
We built this dataset using the Meta-Llama-3-70B-Instruct, and the instruction we used is: Rewrite the question-answer pairs into a paragraph format (Do not use the words 'question' and 'answer' in your responses):.
PMC articles that failed to download are excluded.
Non-medical images (e.g., diagrams) are excluded in an automatic way.
Despite these efforts, this dataset is not… See the full description on the dataset page: https://huggingface.co/datasets/myeongkyunkang/LLaVA-Med-60K-IM-text.LLaVA-Instruct-150K-zh-tw使用openCC將BUAADreamer/llava-en-zh-300k翻譯成繁體中文
coyo-hd-11m-llavanext
Dataset Card for coyo-hd-11m-llavanext
Dataset Summary
This is a data of 22,794,288 synthetic captions for 11,397,144 images from coyo-700m. The "hd" in the title refers to two aspects: high density and high definition. While large alt-text image pair datasets have many images, only a very small proportion of these images are in higher resolutions and have substantial concept density. For example, many of these datasets consist of more than 50% thumbnail sized or very… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/coyo-hd-11m-llavanext.llava-critic-113k
Dataset Card for LLaVA-Critic-113k
🪐 Project Page: https://llava-vl.github.io/blog/2024-10-03-llava-critic/
📰 Paper: https://arxiv.org/abs/2410.02712
🤗 Huggingface Collection: https://huggingface.co/collections/lmms-lab/llava-critic-66fe3ef8c6e586d8435b4af8
👋 Point of Contact: Tianyi Xiong
Dataset Summary
LLaVA-Critic-113k is a high quality critic instruction-following dataset tailored to follow instructions in complex evaluation setting, providing both… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/llava-critic-113k.llava-cc3m-pretrain-595KLLaVA-Med
HealthGPT : A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
Tianwei Lin1, Wenqiao Zhang1, Sijing Li1, Yuqian Yuan1, Binhe Yu2, Haoyuan Li3, Wanggui He3, Hao Jiang3,
Mengze Li4, Xiaohui Song1, Siliang Tang1, Jun Xiao1, Hui Lin1, Yueting Zhuang1, Beng Chin Ooi5
1Zhejiang University,
2University of Electronic Science and Technology of China,
3Alibaba,
4The Hong Kong University of Science and Technology,
5National… See the full description on the dataset page: https://huggingface.co/datasets/acthuan/LLaVA-Med.
