VQA
Datasets
All datasets matching “VQA”VQAv2VQAonline
VQAonline
🌐 Homepage | 🤗 Dataset | 📖 arXiv
Dataset Description
We introduce VQAonline, the first VQA dataset in which all contents originate from an authentic use case.
VQAonline includes 64K visual questions sourced from an online question answering community (i.e., StackExchange).
It differs from prior datasets; examples include that it contains:
(1) authentic context that clarifies the question
(2) an answer the individual asking the question validated as… See the full description on the dataset page: https://huggingface.co/datasets/ChongyanChen/VQAonline.viet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage.
It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English.
The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing,
landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life,
and transportation.Kvasir-VQA-x1
Kvasir-VQA-x1
A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
Kvasir-VQA-x1 on GitHub |
Original Image from Kvasir-VQA(Simula Datasets) |
Paper
🔗 MediaEval Medico 2025 Challenge uses this dataset. We encourage you to check out and participate!
Overview
Kvasir-VQA-x1 is a large-scale dataset designed to benchmark medical visual question answering (MedVQA) in gastrointestinal (GI) endoscopy. It introduces 159,549 new QA… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/Kvasir-VQA-x1.IllusionChar_train
IllusionChar — Training Set
Dataset summary
This repository contains the training split of IllusionChar, the optical character recognition (OCR) component of Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. The task is to transcribe a hidden, case-sensitive alphanumeric sequence from an illusory image, or return No illusion when no sequence is embedded.
Sequences contain 3–5 characters drawn from digits, uppercase Latin letters, and… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionChar_train.VizWiz-VQA
Dataset Card for "VizWiz-VQA"
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of VizWiz-VQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{gurari2018vizwiz,
title={Vizwiz grand… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/VizWiz-VQA.
