datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VQAv2VQAonline
VQAonline
🌐 Homepage | 🤗 Dataset | 📖 arXiv
Dataset Description
We introduce VQAonline, the first VQA dataset in which all contents originate from an authentic use case.
VQAonline includes 64K visual questions sourced from an online question answering community (i.e., StackExchange).
It differs from prior datasets; examples include that it contains:
(1) authentic context that clarifies the question
(2) an answer the individual asking the question validated as… See the full description on the dataset page: https://huggingface.co/datasets/ChongyanChen/VQAonline.Kvasir-VQA-x1
Kvasir-VQA-x1
A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
Kvasir-VQA-x1 on GitHub |
Original Image from Kvasir-VQA(Simula Datasets) |
Paper
🔗 MediaEval Medico 2025 Challenge uses this dataset. We encourage you to check out and participate!
Overview
Kvasir-VQA-x1 is a large-scale dataset designed to benchmark medical visual question answering (MedVQA) in gastrointestinal (GI) endoscopy. It introduces 159,549 new QA… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/Kvasir-VQA-x1.CiQi-VQA
CiQi-Agent
Github | Model | Dataset | Paper
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
Accepted to ECCV 2026
🎯 Overview
CiQi-Agent has been accepted to ECCV 2026.
We present CiQi-Agent, a domain-specific multimodal agent for antique Chinese porcelain connoisseurship. The project is designed to combine fine-grained visual perception, tool-augmented reasoning, and cultural-heritage knowledge… See the full description on the dataset page: https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA.IllusionChar_train
IllusionChar — Training Set
Dataset summary
This repository contains the training split of IllusionChar, the optical character recognition (OCR) component of Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. The task is to transcribe a hidden, case-sensitive alphanumeric sequence from an illusory image, or return No illusion when no sequence is embedded.
Sequences contain 3–5 characters drawn from digits, uppercase Latin letters, and… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionChar_train.VizWiz-VQA
Dataset Card for "VizWiz-VQA"
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of VizWiz-VQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{gurari2018vizwiz,
title={Vizwiz grand… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/VizWiz-VQA.vqa-rad
Dataset Card for VQA-RAD
Dataset Description
VQA-RAD is a dataset of question-answer pairs on radiology images. The dataset is intended to be used for training and testing
Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions.
The dataset is built from MedPix, which is a free open-access online database of medical images.
The question-answer pairs were manually generated by a team of clinicians.… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/vqa-rad.IllusionChar_test
IllusionChar — Test Set
Dataset summary
This repository contains the public test split of IllusionChar, the OCR benchmark introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each metadata row can be paired across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control conditions.
The expected output for an illusion-bearing or source-condition image is an exact, case-sensitive… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionChar_test.path-vqa
Dataset Card for PathVQA
Dataset Description
PathVQA is a dataset of question-answer pairs on pathology images. The dataset is intended to be used for training and testing
Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions.
The dataset is built from two publicly-available pathology textbooks: "Textbook of Pathology" and "Basic Pathology", and a
publicly-available digital library: "Pathology… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/path-vqa.OK-VQAMNIST_train
IllusionMNIST — Training Set
Dataset summary
This repository contains the training split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. The dataset is intended for training models to recognize MNIST digits embedded as visual illusions (pareidolia) in generated scenes and to reject images that contain no illusion.
MNIST source-condition images were sampled and resized to 512 × 512 pixels, combined… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_train.FashionMnist_train
IllusionFashionMNIST — Training Set
Dataset summary
This repository contains the training split of IllusionFashionMNIST, one of the four datasets introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. It is designed to train and evaluate models on the recognition of Fashion-MNIST categories embedded as visual illusions (pareidolia) in generated scenes.
The source-condition images are sampled from Fashion-MNIST and resized to… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/FashionMnist_train.Awesome_Spatial_VQA_BenchmarksIllusionAnimals_train
IllusionAnimals — Training Set
Dataset summary
This repository contains the training split of IllusionAnimals, one of the four benchmarks introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. It supports training models to identify animal categories embedded as visual illusions (pareidolia) in generated scenes and to recognize when no illusion is present.
The source-condition animal images were generated with SDXL-Lightning.… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionAnimals_train.FashionMnist_test
IllusionFashionMNIST — Test Set
Dataset summary
This repository contains the public test split of IllusionFashionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each metadata row identifies a Fashion-MNIST target and can be paired across five image conditions: source-condition, illusion, filtered illusion, illusionless control, and filtered illusionless control.
The source-condition images originate from… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/FashionMnist_test.MNIST_test
IllusionMNIST — Test Set
Dataset summary
This repository contains the public test split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Every indexed example can be compared across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control images.
The source-condition images are sampled from MNIST and resized to 512 × 512 pixels. Illusion images were… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_test.OCR-VQA
Dataset Card for "OCR-VQA"
More Information needed
IllusionAnimals_test
IllusionAnimals — Test Set
Dataset summary
This repository contains the public test split of IllusionAnimals, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each annotated example is paired across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control conditions.
The animal source-condition images were generated with SDXL-Lightning. English scene descriptions and… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionAnimals_test.DeepTumorVQA_2.0
DeepTumorVQA v2
3D abdominal-CT diagnostic Visual Question Answering benchmark with 42
clinical subtypes and 438K total QA pairs (10K curated benchmark + 428K
training pool). Includes pre-extracted 2D and video modalities, 20K agent
training trajectories with tool-use traces, and a paper-locked leaderboard.
Resources
📄 Paper (arXiv)
https://arxiv.org/abs/2605.09679
💻 Code (GitHub)
https://github.com/Schuture/DeepTumorVQA
🤗 Dataset (this… See the full description on the dataset page: https://huggingface.co/datasets/tumor-vqa/DeepTumorVQA_2.0.vqav2-smallPMC-VQA
PMC-VQA Dataset
PMC-VQA Dataset
Daraset Structure
Sample
Dataset Structure
PMC-VQA (version-1: 227k VQA pairs of 149k images).
train.csv: metafile of train set
test.csv: metafile of test set
test_clean.csv: metafile of test clean set
images.zip: images folder
(update version-2: noncompound images).
train2.csv: metafile of train set
test2.csv: metafile of test set
images2.zip: images folder
Sample
A row in train.csv is shown bellow… See the full description on the dataset page: https://huggingface.co/datasets/RadGenome/PMC-VQA.Japanese-Medical-VQA-12m
Japanese Medical VQA 12M
Japanese Medical VQA 12M is a large-scale Japanese medical multimodal dataset built from Open-PMC-18M and released in Parquet and Webdataset format.
This dataset contains outputs from multiple data-construction stages, including:
source captions
Japanese translations of source captions
enriched captions
Japanese translations of enriched captions
question-answering
Current Repository Format
This repository currently stores the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/MIL-UT/Japanese-Medical-VQA-12m.VQAv2_train
Dataset Card for "VQAv2_train"
More Information needed
MedPix-VQA
MedPix-VQA Dataset
The MedPix-VQA dataset is a version of the data found at MEDPIX-ClinQA, specifically modified to address an image overlap issue that would result from directl splitting the original dataset. This overlap can lead to a model potentially seeing the same image during both training and validation, potentially leading to bias or data leakage.
Key Modifications:
We have modified the dataset to ensure no image overlap between the training and validation… See the full description on the dataset page: https://huggingface.co/datasets/mmoukouba/MedPix-VQA.vqa
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
This version includes all images in the dataset. For a more lightweight and accessible alternative, please refer to the (1.1 release)[https://huggingface.co/datasets/worldcuisines/vqa-v1.1/] which reduces download size while preserving all text and metadata.
The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆.
WorldCuisines is a… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa.SLAKE-vqa-english
Dataset Description
SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering [ISBI 2021 oral]
Corresponding Authors: Bo Liu, Xiao-Ming Wu
Original dataset is retrieved from https://huggingface.co/datasets/BoKelvin/SLAKE
In this dataset, we modified some things to match our task:
The original dataset are bilingual, we filtered to only English
We only take the image (converted as PIL object), question, and answer column
Any questions, please… See the full description on the dataset page: https://huggingface.co/datasets/mdwiratathya/SLAKE-vqa-english.CT-RATE-VQA
CT-RATE-VQA Dataset
We constructed a large-scale CT-VQA dataset based on the ReXGroundingCT data \cite{rexct} to support model training and evaluation.For each case, CT volumes were processed along with their corresponding multi-class segmentation masks, where each mask channel represents a specific lesion type.
This dataset is designed for medical visual question answering (Med-VQA) tasks.
PMC-Clinical-VQA
PMC-VQA: A Large-Scale Visual Question Answering Dataset for Clinical Figures
This dataset contains over 1,700,000 Visual Question Answering (VQA) samples derived from figures and charts in biomedical articles from PubMed Central (PMC).
This is a preliminary release. A full dataset card and an accompanying research paper are currently in preparation.
Raw version of this dataset with licenses and metadata can be found on Hugging Face: DermaVLM/pmc_clinical_VQA_raw
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DermaVLM/PMC-Clinical-VQA.vqasynth_spacellava
VQASynth_spacellava
Uses the VQASynth pipeline to synthesize spatialVQA samples, mixed with general VQA samples used to fine-tune LLaVA-v1.5-13b.
RadImageNet-VQA
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
We introduce RadImageNet-VQA, a large-scale dataset designed for training and benchmarking radiologic VQA on CT and MRI exams. Built from the CT/MRI subset of RadImageNet and its expert-curated anatomical and pathological annotations, RadImageNet-VQA provides 750K images with 7.5M generated samples, including 750K medical captions for visual-text alignment and 6.75M… See the full description on the dataset page: https://huggingface.co/datasets/raidium/RadImageNet-VQA.
