datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VQAonline
VQAonline
🌐 Homepage | 🤗 Dataset | 📖 arXiv
Dataset Description
We introduce VQAonline, the first VQA dataset in which all contents originate from an authentic use case.
VQAonline includes 64K visual questions sourced from an online question answering community (i.e., StackExchange).
It differs from prior datasets; examples include that it contains:
(1) authentic context that clarifies the question
(2) an answer the individual asking the question validated as… See the full description on the dataset page: https://huggingface.co/datasets/ChongyanChen/VQAonline.SuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.vqa
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
This version includes all images in the dataset. For a more lightweight and accessible alternative, please refer to the (1.1 release)[https://huggingface.co/datasets/worldcuisines/vqa-v1.1/] which reduces download size while preserving all text and metadata.
The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆.
WorldCuisines is a… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa.LRS-VQA
LRS-VQA Dataset
This repository contains the LRS-VQA benchmark dataset, presented in the paper When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning.
Code: The associated code and evaluation scripts can be found on the project's GitHub repository: https://github.com/VisionXLab/LRS-VQA
Introduction
Efficient vision-language understanding of large Remote Sensing Images (RSIs) is meaningful but challenging. Current… See the full description on the dataset page: https://huggingface.co/datasets/ll-13/LRS-VQA.viet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.medical-vqaTraffic-VQAvqa-v1.1
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
This version removes all images from the 1.0 release to reduce download size and improve accessibility. All text and metadata remain unchanged.
The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆.
WorldCuisines is a massive-scale visual question answering (VQA) benchmark for multilingual and multicultural understanding through… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa-v1.1.vqav2-full-metadataEG-VQA
EG-VQA
This repository provides the videos and annotations of EG-VQA. EG-VQA is an evidence-grounded, open-ended Video Question Answering benchmark. It contains 2,067 videos and 11,838 question-answer pairs. Each question is paired with one or more temporally localized evidence segments, including timestamps and textual descriptions.
Questions cover four types:
Descriptive: recognizing and describing visible events or states.
Temporal: reasoning about event order, timing… See the full description on the dataset page: https://huggingface.co/datasets/lphuang33/EG-VQA.GAP-VQA-Datasetnpnp-nonogram-vqavqa-cmsv-benchmark
VQA-CMSV Benchmark Data Package
This repository contains annotation splits for VQA v2-CMSV, GQA-CMSV, and VG-CMSV, plus patch-mask NPZ files used for mask supervision experiments.
Contents
data/vqa_v2_cmsv/train.json, data/vqa_v2_cmsv/val.json, data/vqa_v2_cmsv/test.json
data/gqa_cmsv/train.jsonl, data/gqa_cmsv/val.jsonl, data/gqa_cmsv/test.jsonl
data/vg_cmsv/train.jsonl, data/vg_cmsv/val.jsonl, data/vg_cmsv/test.jsonl
masks/vqa_v2_cmsv_masks.npz
masks/gqa_cmsv_masks.npz… See the full description on the dataset page: https://huggingface.co/datasets/as-benchmark-artifacts/vqa-cmsv-benchmark.YTTB-VQA
Dataset Card for Dataset Name
Dataset Summary
The YTTB-VQA Dataset is a collection of 400 Youtube thumbnail question-answer pairs to evaluate the visual perception abilities of in-text images. It covers 11
categories, including technology, sports, entertainment, food, news, history, music, nature, cars, and education.
Supported Tasks and Leaderboards
This dataset supports many tasks, including visual question answering, image captioning, etc.
License… See the full description on the dataset page: https://huggingface.co/datasets/mlpc-lab/YTTB-VQA.PMC-VQA-text
PMC-VQA-text
This dataset is a text format of PMC-VQA.
We built this dataset using the Meta-Llama-3-70B-Instruct, and the instruction we used is: Rewrite the question-answer pairs into a paragraph format (Do not use the words 'question' and 'answer' in your responses):.
train_text.json corresponds to the train.csv and train_2.csv splits in the PMC-VQA dataset.
Samples with two or more question-and-answer pairs were selected.
Citation
If you find this dataset useful… See the full description on the dataset page: https://huggingface.co/datasets/myeongkyunkang/PMC-VQA-text.spatial_mosaic_vqa
SpatialMosaic: A Multi-View VLM Dataset for Partial Visibility
Description
SpatialMosaic is a multi-view visual question answering dataset for evaluating spatial reasoning under partial visibility, occlusion, and low-overlap views. It pairs indoor ScanNet++ and outdoor Waymo scene references with multi-frame VQA annotations. Questions require models to combine fragmented evidence across 2-5 views, rather than answering from a single image. The tasks… See the full description on the dataset page: https://huggingface.co/datasets/jmkey/spatial_mosaic_vqa.VinDR-CXR-VQA
VinDr-CXR-VQA Dataset
Dataset Description
VinDr-CXR-VQA is a large-scale chest X-ray Visual Question Answering (VQA) dataset designed for explainable medical AI with spatial grounding capabilities. The dataset combines natural language question-answer pairs with bounding box annotations and clinical reasoning explanations.
Key Features
🏥 4,394 chest X-ray images from VinDr-CXR
💬 17,597 question-answer pairs across 6 question types
📍 Spatial… See the full description on the dataset page: https://huggingface.co/datasets/Dangindev/VinDR-CXR-VQA.VQA-RADvietnamese-traffic-sign-vqa
Vietnamese Traffic Sign VQA
Visual Question Answering dataset for Vietnamese traffic signs.
Built from Kaggle VNTS (CC BY-SA 4.0).
Statistics
Split
Images
QA Pairs
QA/Image
Train
2,193
104,146
47.5
Val
272
12,944
47.6
Test
271
12,966
47.8
Total
2,736
130,056
47.5
Question Types
12 types: yes_no, count, sign_type, color, shape, location, attribute, negative, spatial_rel, count_total, multi_object, context
Format
{… See the full description on the dataset page: https://huggingface.co/datasets/Anakonkai/vietnamese-traffic-sign-vqa.omni3d_857_3dg_vqa_capVinDR-CXR-VQA
VinDr-CXR-VQA Dataset
Dataset Description
VinDr-CXR-VQA is a large-scale chest X-ray Visual Question Answering (VQA) dataset designed for explainable medical AI with spatial grounding capabilities. The dataset combines natural language question-answer pairs with bounding box annotations and clinical reasoning explanations.
Key Features
🏥 4,394 chest X-ray images from VinDr-CXR
💬 17,597 question-answer pairs across 6 question types
📍 Spatial grounding with… See the full description on the dataset page: https://huggingface.co/datasets/faizan711/VinDR-CXR-VQA.unsloth-pmc-vqa-trsurgvu-cat2-vqa
SurgVU Category 2 VQA pairs
No video or frames are included. This is 23,354 question–answer pairs over
30-second windows of the published SurgVU dataset. Each record gives a case id and
a start/stop time, so anyone with the SurgVU videos can regenerate the exact frames.
Used to train the vision-language model in our SurgVU 2026 Category 2 submission.
Contents
file
data/train.jsonl
18,619 pairs
data/val.jsonl
4,735 pairs
recipe/build_qa_pairs.py… See the full description on the dataset page: https://huggingface.co/datasets/opscribe-ai/surgvu-cat2-vqa.turkish-medical-vqa-evaluatedPMC-VQAPATH-VQAM3D-VQAUnLOK-VQA
📊 Dataset: UnLOK-VQA (Unlearning Outside Knowledge VQA)
Paper: Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
Code: https://github.com/Vaidehi99/mmmedit
Link: Dataset Link
This dataset contains approximately 500 entries with the following key attributes:
"id": Unique Identifier for each entry
"src": The question whose answer is to be deleted ❓
"pred": The answer to the question meant for deletion ❌
"loc": Related neighborhood questions… See the full description on the dataset page: https://huggingface.co/datasets/vaidehi99/UnLOK-VQA.CS381V-hardest-vqamedical-vqa-robustness-analysis
Medical VQA Robustness Analysis
This dataset contains robustness analysis results for medical vision-language models on chest X-ray visual question answering tasks. The analysis evaluates model performance under various question perturbations to assess clinical safety and response stability.
Source Data
This analysis is based on the MIMIC-CXR-VQA dataset from PhysioNet, which provides chest X-ray images paired with clinically relevant questions and answers.
Models… See the full description on the dataset page: https://huggingface.co/datasets/saillab/medical-vqa-robustness-analysis.
