datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VisualWebInstruct
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs.
Please also checkout our more recent verified version at Huggingface.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct.VisualWebInstruct-Recall
Introduction
This is the dataset recalled from Google Search from the seed images.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
VisualWebInstruct-verified
🧠 VisualWebInstruct-Verified: High-Confidence Multimodal QA for Reinforcement Learning
VisualWebInstruct-Verified is a high-confidence subset of VisualWebInstruct, curated specifically for Reinforcement Learning (RL) and Reward Model training.
It contains verified multimodal question–answer pairs where correctness, reasoning quality, and image–text alignment have been explicitly validated.
This dataset is ideal for RLVR training pipelines.
📘 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct-verified.VisualWebInstruct-Seed
Introduction
This is the seed dataset we used to conduct Google Search.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
VisualWebInstruct
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs.
Links
GitHub Repository
Research Paper
Project Website… See the full description on the dataset page: https://huggingface.co/datasets/taoye1992/VisualWebInstruct.VisualReasoner-30k
Dataset Card for VisualReasoner-30k
Dataset Details
This dataset is an extension of VisualReasoner-1M, containing approximately 30k cases and can be used for training visual reasoning tasks.
Unlike VisualReasoner-1M, this dataset models the reasoning process in an end-to-end format to better accommodate scenarios where explicit tool invocation is not allowed.
Dataset Descriptions
The structure of each case is as follows:
{
"identity": "Case ID"… See the full description on the dataset page: https://huggingface.co/datasets/orange-sk/VisualReasoner-30k.Turkish-medical-visual-question-answering-LLaVa-dataset
Türkçe Radyoloji Görüntüleme Veri Seti - data_RAD
data_RAD veri seti, radyoloji görüntüleri üzerinde görsel soru-cevaplama (VQA) araştırmaları yapmak amacıyla Türkçeye çevrilmiş ve LLaVa mimarisiyle uyumlu hale getirilmiştir. Bu veri seti, tıbbi görüntü analizi ve yapay zeka destekli radyoloji uygulamalarını geliştirmek için kullanılabilir.
Veri Seti İçeriği
Toplam Görüntü Sayısı: 316
Veri Yapısı: DatasetDict({ train: Dataset({ features: ['image'], num_rows: 316 }) })
Özellikler:… See the full description on the dataset page: https://huggingface.co/datasets/nezahatkorkmaz/Turkish-medical-visual-question-answering-LLaVa-dataset.Visual-Math-Eval
Visual Equation Solving Benchmark
This repository contains the dataset introduced in the paper:
Can Vision-Language Models Solve Visual Math Equations? which is currently accepted in EMNLP 2025 (Main)
Despite strong performance in vision and language understanding, Vision-Language Models (VLMs) struggle on tasks requiring integrated perception and symbolic reasoning. This benchmark evaluates VLMs on visual equation solving, where systems of linear equations are represented using… See the full description on the dataset page: https://huggingface.co/datasets/monjoychoudhury29/Visual-Math-Eval.VisualMRC
VisualMRC
VisualMRC: Machine Reading Comprehension on Document Images
📖 arXiv 🌐 github
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
Citation and contact
If you use this dataset, please cite our work:
@inproceedings{VisualMRC2021,
author = {Ryota Tanaka and
Kyosuke Nishida and
Sen Yoshida},
title =… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/VisualMRC.VisualReasoner-1M
Dataset Card for VisualReasoner-1M
Dataset Details
This is a dataset for the paper From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis. The dataset contains approximately 1 million cases and can be used for training visual reasoning tasks. The reasoning process involves breaking down tasks and utilizing tools to solve complex and challenging visual question-answering tasks progressively.
For detailed data synthesis methods, please… See the full description on the dataset page: https://huggingface.co/datasets/orange-sk/VisualReasoner-1M.VisualReferPrompt
Special Note
We have open-sourced a preliminary version of our dataset.
However, please note that the experimental version of the dataset, which includes additional images,
is currently undergoing review by our school's ethics committee.
We will update the repository with the latest version of the dataset as soon as possible. Thank you for your understanding.
Dataset Card for Dataset Name
vrpbench is a benchmark dataset designed for visual referring… See the full description on the dataset page: https://huggingface.co/datasets/zongjieli/VisualReferPrompt.
