datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Leopard-Instruct
Leopard-Instruct
Paper | Github | Models-LLaVA | Models-Idefics2
Summaries
Leopard-Instruct is a large instruction-tuning dataset, comprising 925K instances, with 739K specifically designed for text-rich, multiimage scenarios. It's been used to train Leopard-LLaVA [checkpoint] and Leopard-Idefics2 [checkpoint].
Loading dataset
to load the dataset without automatically downloading and process the images (Please run the following codes with datasets==2.18.0)… See the full description on the dataset page: https://huggingface.co/datasets/wyu1/Leopard-Instruct.LLaVA-OneVision-1.5-Instruct-Data
LLaVA-OneVision-1.5 Instruction Data
Paper | Code
📌 Introduction
This dataset, LLaVA-OneVision-1.5-Instruct, was collected and integrated during the development of LLaVA-OneVision-1.5. LLaVA-OneVision-1.5 is a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. This meticulously curated 22M instruction dataset (LLaVA-OneVision-1.5-Instruct) is part of a… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Instruct-Data.Innovator-VL-Instruct-46M
Innovator-VL-Instruct-46M
Paper | Code
🤗🤗 The data is being uploaded continuously
Introduction
To further enhance the model’s ability to handle a broad range of visual tasks with accurate, grounded, and instruction-aligned responses, we perform full-parameter visual instruction supervised fine-tuning (SFT).This SFT stage serves as a critical bridge between multimodal pretraining and subsequent reinforcement learning, providing both general capability coverage and a… See the full description on the dataset page: https://huggingface.co/datasets/InnovatorLab/Innovator-VL-Instruct-46M.instructpix2pix-clip-filtered
Dataset Card for InstructPix2Pix CLIP-filtered
Dataset Summary
The dataset can be used to train models to follow edit instructions. Edit instructions
are available in the edit_prompt. original_image can be used with the edit_prompt and
edited_image denotes the image after applying the edit_prompt on the original_image.
Refer to the GitHub repository to know more about
how this dataset can be used to train a model that can follow instructions.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/timbrooks/instructpix2pix-clip-filtered.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.MTabVQA-InstructPaper
MTabVQA-Instruct Sub-datasets
This directory contains multiple MTabVQA-Instruct datasets for visual question answering over tables.
Datasets
MTabVQA-Atis-Instruct
MTabVQA-MiMo-Instruct
MTabVQA-Multitab-Instruct
MTabVQA-Spider-Instruct
Each dataset contains a VQA.jsonl file and a table_images directory with the corresponding table images.
Important Note for Multitab-Instruct
You must unzip the table_images.zip file in MTabVQA-Multitab-Instruct/ to access… See the full description on the dataset page: https://huggingface.co/datasets/mtabvqa/MTabVQA-Instruct.llava-instruct-mix
LLaVA Instruct Mix
Added OCR and Chart QA dataset into this for more text extraction questions
InstructPart
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
Zifu Wan, Yaqi Xie, Ce Zhang, Zhiqiu Lin, Zihan Wang, Simon Stepputtis, Deva Ramanan, Katia SycaraRobotics Institute, Carnegie Mellon UniversityACL 2025 Main
👀 Introduction
This dataset contains annotated images and language instructions for task-oriented part segmentation. The dataset is divided into three main folders: all, train, and test.
📦 Download
Goodle Drive
⏳ Folder… See the full description on the dataset page: https://huggingface.co/datasets/zifuwan/InstructPart.TS_Instruct_Reasoning_V1Chat-TS training datasets
This dataset is intended to train reasoning.
The dataset consists of real-world time-series with synthetic text.
If you use this work please cite:
@misc{quinlan2025chattsenhancingmultimodalreasoning,
title={Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and Natural Language Data},
author={Paul Quinlan and Qingguo Li and Xiaodan Zhu},
year={2025},
eprint={2503.10883},
archivePrefix={arXiv},
primaryClass={cs.AI}… See the full description on the dataset page: https://huggingface.co/datasets/PaulQ1/TS_Instruct_Reasoning_V1.llava-instruct-mix-vsfttheblackcat102/llava-instruct-mix reformated for VSFT with TRL's SFT Trainer.
See https://github.com/huggingface/trl/blob/main/examples/scripts/vsft_llava.py.
MetaQuery_Instruct_2.4M_512resThe data is licensed CC-by-NC. Third party content pulled from other locations are subject to their own licenses and you may have other legal obligations or restrictions that govern your use of that content.
The MetaQuery dataset is also released under ODC-BY and Common Crawl terms of use, because it is sourced from mmc4.
docatlas_instruct
DocAtlas: Multilingual Document Understanding Across 80+ Languages
DocAtlas is a large-scale, high-fidelity multilingual OCR dataset covering 82 languages and 10 writing systems, built through model-free differential rendering. It provides precise structural annotations in a unified DocTag format encoding layout, text, and component types — without relying on any learned models for core annotation.This dataset powers the training of DocAtlas-DeepSeek, which achieves… See the full description on the dataset page: https://huggingface.co/datasets/ahmedheakl/docatlas_instruct.food-visual-instructions
Adapting Multimodal Large Language Models to Domains via Post-Training (EMNLP 2025)
This repos contains the food visual instructions for post-training MLLMs in our paper: On Domain-Specific Post-Training for Multimodal Large Language Models.
The main project page is: Adapt-MLLM-to-Domains
Data Information
Using our visual instruction synthesizer, we generate visual instruction tasks based on the image-caption pairs from extended Recipe1M+ dataset. These synthetic… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/food-visual-instructions.LLaVA-Instruct-150K-zh-tw使用openCC將BUAADreamer/llava-en-zh-300k翻譯成繁體中文
KoLLaVA-v1.5-Instruct-581k
KoLLaVA-v1.5-Instruct-581k
한국어 Vision-Language 모델을 위한 instruction tuning 데이터셋입니다.
데이터셋 정보
총 샘플 수: 435,093개
형식: ChatML 형식 (role: user/assistant, content: 텍스트)
이미지: COCO + GQA + Visual Genome 데이터셋
언어: 한국어
포함된 데이터셋
COCO 데이터: 362,953개 샘플
MS COCO 2017 이미지 기반
한국어 대화 데이터
GQA 데이터: 72,140개 샘플
GQA (Visual Question Answering) 이미지 기반
한국어 대화 데이터
Visual Genome 데이터: 포함
Visual Genome 이미지 기반
한국어 대화 데이터
제외된 데이터셋
EKVQA 데이터: AI Hub 라이선스로 인해 공개 불가… See the full description on the dataset page: https://huggingface.co/datasets/ko-vlm/KoLLaVA-v1.5-Instruct-581k.LLaVA-Instruct-600K-Chinese
仿照 LLaVA-Instruct-150K ,使用 Qwen2.5-VL-32B-Instruct 合成的用于微调中文VLM的数据;也可以与英文数据集混合使用,训练多语言VLM
任务类型为基于单张图片的问答和对话,每个样本都对应一张不同的图片,其中大部分图片包含中文字符,更适合中文场景下视觉语言模型的训练。
图片从各类中文网站上爬取
包含3类任务:日常对话、复杂推理、描述图片。日常对话通常是5轮对话,其余任务是1轮对话。
每种任务的数量如下:
任务类型
数量
日常对话
247,431
复杂推理
194,646
描述图片
199,791
用于生成对话数据的prompt如下
日常对话
设计一个你和一个询问这张照片的人之间的对话。答案应该是视觉AI助手看到图像并回答问题的语气。
你需要提出不同的问题并给出相应的答案。问题可以包括询问图像视觉内容的问题,包括对象类型、对象计数、对象动作、对象位置、对象之间的相对位置等。必须是有明确答案的问题,即
(1) 人们可以在图像中明确看到问题所问的内容,并且可以自信地回答;
(2)… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/LLaVA-Instruct-600K-Chinese.MathCanvas-Instruct
MathCanvas-Instruct Dataset
🚀 Data Usage
from datasets import load_dataset
dataset = load_dataset("shiwk24/MathCanvas-Instruct")
print(dataset)
📖 Overview
MathCanvas-Instruct is a high-quality, fine-tuning dataset with 219K examples of interleaved visual-textual reasoning paths. It is the core component for the second phase of the [MathCanvas] framework: Strategic Visual-Aided Reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/shiwk24/MathCanvas-Instruct.llava-instruct-mix
LLaVA Instruct Mix
Summary
The LLaVA Instruct Mix dataset is a processed version of LLaVA Instruct Mix.
Data Structure
Format: Conversational
Type: Language-modeling
Columns:
"images": The image associated with the text.
"prompt": A list of messages that form the context for the conversation.
"completion": The last message in the conversation, which is the model's response.
This structure allows models to learn from the context of the conversation… See the full description on the dataset page: https://huggingface.co/datasets/trl-lib/llava-instruct-mix.MetaQuery_Instruct_2.4MThe data is licensed CC-by-NC. Third party content pulled from other locations are subject to their own licenses and you may have other legal obligations or restrictions that govern your use of that content.
The MetaQuery dataset is also released under ODC-BY and Common Crawl terms of use, because it is sourced from mmc4.
starcoder2-instruct-assetsinstructpix2pix-10-samples
Dataset Card for "test"
More Information needed
llava-instruct-mix-vsft-miniOriginally from https://huggingface.co/datasets/HuggingFaceH4/llava-instruct-mix-vsft but 0.33% randomnly sampled
llava-1.5-665k-instructionsThis dataset repository, LLaVA-1.5-665K-Instructions, is notably utilized in the paper Zero-Shot Vision Encoder Grafting via LLM Surrogates.
The official code repository for the paper can be found here: https://github.com/kaiyuyue/zero
LLaVA-1.5-665K-Instructions
This dataset repo contains the entire LLaVA-1.5-665K-Instructions in one place, including images and text sequences.
The images are in train_split/*.tars and the text sequences are in jsons:
llava_v1_5_mix665k.json is the… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/llava-1.5-665k-instructions.MMc-Instruct-Stage2RoboTwin_instruct-pix2pix
数据集介绍
Contributor: Shenghao Yang, Yimou Wu
数据简介
本数据集是通过https://github.com/TianxingChen/RoboTwin [1]单臂机器人模拟器在 block_hammer_beat, block_handover, blocks_stack_easy 三个任务上间隔50帧采样得到的,其数据规模如下:
block_hammer_beat block_handover blocks_stack_easy
train 100 100 100
val 10 10 10
test 10 10 10
整体数据规模如下:
train val test
num 300 30 30… See the full description on the dataset page: https://huggingface.co/datasets/YimouWu/RoboTwin_instruct-pix2pix.instructpix2pix_toontown_100einstructverse_mask_v1.0instructpix2pix-clip-filtered-upscaledTS_Instruct_QA_GoldWelcome to the TS_Instruct_QA_Gold dataset
This dataset is intended to evaluate time-series reasoning.
The dataset consists of real-world time-series with synthetic text.
The dataset was human evaluated for correctness
@misc{quinlan2025chattsenhancingmultimodalreasoning,
title={Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and Natural Language Data},
author={Paul Quinlan and Qingguo Li and Xiaodan Zhu},
year={2025},
eprint={2503.10883}… See the full description on the dataset page: https://huggingface.co/datasets/PaulQ1/TS_Instruct_QA_Gold.wordocr_instruct_v4
