datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset_viahe_vqa
Vietnamese Sidewalk Violation VQA (vi phạm vỉa hè)
Bộ dữ liệu Hỏi–Đáp trên ảnh (VQA) tiếng Việt đầu tiên về hành vi lấn chiếm/vi phạm
sử dụng vỉa hè, gắn với căn cứ pháp lý Nghị định 168/2024/NĐ-CP. Xây dựng cho đồ án
môn học SE365 (Trường ĐH Công nghệ Thông tin, ĐHQG-HCM).
Ảnh: 6.349 ảnh
Cặp hỏi–đáp: 19.474
Ngôn ngữ: Tiếng Việt
Câu hỏi: 10 câu cố định, 4 loại (yes/no, what đa nhãn, đếm số lượng, không gian)
IAA (2 người gán nhãn độc lập): macro-κ tăng 0,736 → 0,854 qua 3 vòng… See the full description on the dataset page: https://huggingface.co/datasets/manhdungcr7/dataset_viahe_vqa.M-VQA
M-VQA: Visual Question Answering under Image Distortions
Dataset Description
M-VQA is the official dataset of the M-VQA Challenge, an ACM Multimedia 2026 Grand Challenge on evaluating visual question answering systems under image distortions.
The dataset contains 12,400 image-question-answer samples:
400 original images;
12,000 distorted images generated from the original images using 30 distortion types;
five distortion severity levels, numbered from 4 to 8.… See the full description on the dataset page: https://huggingface.co/datasets/AIBench/M-VQA.vqa
SupraBench: Molecular Identification (VQA)
Auxiliary vision task of SupraBench: given
the 2D depiction of a supramolecular host or guest, identify the molecule (its
name / alias set, and where available its canonical SMILES). It probes whether
multimodal LLMs can ground chemical structure from images alone.
Links
📄 Paper: arXiv:2606.13477
💻 Code: https://github.com/Tianyi-Billy-Ma/SupraBench
🤗 All datasets: https://huggingface.co/SupraBench… See the full description on the dataset page: https://huggingface.co/datasets/SupraBench/vqa.multilingual-vqacamerabench_vqa_lmms_evalJIC-VQA
JIC-VQA
Dataset Description
Japanese Image Classification Visual Question Answering (JIC-VQA) is a benchmark for evaluating Japanese Vision-Language Models (VLMs). We built this benchmark based on the recruit-jp/japanese-image-classification-evaluation-dataset by adding questions to each sample. All questions are multiple-choice, each with four options. We select options that closely relate to their respective labels in order to increase the task's difficulty.
The… See the full description on the dataset page: https://huggingface.co/datasets/line-corporation/JIC-VQA.vqa_radLicense: Apache-2.0
This dataset is derived from flaviagiammarino/vqa-rad.All credit goes to the original authors.
The dataset has been reformatted into .tsv files for compatibility with VLMEvalKit.No changes were made to the original image or text content.
This version is intended for benchmarking vision-language models using VLMEvalKit.
RadFig-VQA
RadFig-VQA Dataset
Overview
RadFig-VQA is a large-scale medical visual question answering dataset based on radiological figures from PubMed Central (PMC). The dataset comprises 70,550 images and 238,294 question-answer pairs - making it the largest radiology-specific VQA dataset by number of QA pairs - generated from radiological figures across diverse imaging modalities and clinical contexts, designed for comprehensive medical VQA evaluation.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/YYama0/RadFig-VQA.MMVMBench_VQAGenerative-VQA-V2-Curated
Generative-VQA-V2-Curated
A curated, balanced, and cleaned version of the VQA v2 dataset specifically optimized for Generative Visual Question Answering.
This dataset transforms the standard VQA task into a generative challenge by removing "yes/no" shortcuts and balancing answer distributions to prevent model over-fitting on dominant classes.
Dataset Summary
The primary goal of this curated set is to provide a "clean" signal for training multimodal models by:
Eliminating… See the full description on the dataset page: https://huggingface.co/datasets/Deva8/Generative-VQA-V2-Curated.afrimed-vqa-parallel
AfriMed-VQA-Parallel
AfriMed-VQA-Parallel is a multi-way parallel visual question answering benchmark covering 14 languages.
Requesting Access
This dataset is gated for research tracking. Please fill in the access request form above to request access.
C-VQAThe dataset repo contains the data for C-VQA-Real dataset, for complete data and evaluating your model on our dataset, please refer to https://github.com/Letian2003/C-VQA.
VQA_SUNRGBD_v23dct_lung_vqa_with_reasoningEL_VQAMedJourneyBench_Multi_Image_VQA
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/JourneyBench/JourneyBench_Multi_Image_VQA.AVE-Assessing_VQA_Evaluators-v1Samples with sample_id from 38 to 7421 serve as the test set, while others are for validation. The division barely introduces shift on the distribution across source VQA datasets.
This is the official repository (still under construction) for paper "Towards Flexible Evaluation for Generative Visual Question Answering", an oral in ACM Multimedia 2024.
Please refer to our github repository (https://github.com/jihuishan/flexible_evaluation_for_vqa_mm24/tree/main) for more details.
Welcome any… See the full description on the dataset page: https://huggingface.co/datasets/Huishan/AVE-Assessing_VQA_Evaluators-v1.IC-VQAdefuser-vqa-oe_results3dct_better_vqaVQA-phi3VQAv2Val2014VQAv3expert-vqa-no-manual_resultssurgery-vqa-benchmark
Surgery VQA Benchmark
手术视频问答基准数据集
数据集
LensID
lensid_lens_train.tsv / lensid_lens_test.tsv - 晶状体分割 VQA
lensid_pupil_train.tsv / lensid_pupil_test.tsv - 瞳孔分割 VQA
lensid_phase_train.tsv / lensid_phase_test.tsv - 手术阶段识别 VQA
CatRelDet
catreldet_train.tsv / catreldet_test.tsv - 白内障手术相关性检测 VQA
格式
TSV 文件包含以下字段:
index: 唯一标识符
question: 问题
image: Base64 编码的图像
category: 任务类别
answer: 正确答案 (A/B/C/D)
A, B, C, D: 四个选项
expert-vqa_resultsnortheast_vqavibook-OCR_VQAdefuser-vqa-mcq_resultsVQAv2_ko
