datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VQAonline
VQAonline
🌐 Homepage | 🤗 Dataset | 📖 arXiv
Dataset Description
We introduce VQAonline, the first VQA dataset in which all contents originate from an authentic use case.
VQAonline includes 64K visual questions sourced from an online question answering community (i.e., StackExchange).
It differs from prior datasets; examples include that it contains:
(1) authentic context that clarifies the question
(2) an answer the individual asking the question validated as… See the full description on the dataset page: https://huggingface.co/datasets/ChongyanChen/VQAonline.vqa
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
This version includes all images in the dataset. For a more lightweight and accessible alternative, please refer to the (1.1 release)[https://huggingface.co/datasets/worldcuisines/vqa-v1.1/] which reduces download size while preserving all text and metadata.
The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆.
WorldCuisines is a… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa.viet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.medical-vqaTraffic-VQAvqa-v1.1
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
This version removes all images from the 1.0 release to reduce download size and improve accessibility. All text and metadata remain unchanged.
The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆.
WorldCuisines is a massive-scale visual question answering (VQA) benchmark for multilingual and multicultural understanding through… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa-v1.1.GAP-VQA-DatasetPMC-VQA-text
PMC-VQA-text
This dataset is a text format of PMC-VQA.
We built this dataset using the Meta-Llama-3-70B-Instruct, and the instruction we used is: Rewrite the question-answer pairs into a paragraph format (Do not use the words 'question' and 'answer' in your responses):.
train_text.json corresponds to the train.csv and train_2.csv splits in the PMC-VQA dataset.
Samples with two or more question-and-answer pairs were selected.
Citation
If you find this dataset useful… See the full description on the dataset page: https://huggingface.co/datasets/myeongkyunkang/PMC-VQA-text.VQA-RADPMC-VQAturkish-medical-vqa-evaluatedtdtu_vqa_dataset_herb
TDTU VQA Dataset — Vietnamese Medicinal Herbs 🌿
Dataset Description
TDTU VQA Dataset Herb is a Vietnamese Visual Question Answering (VQA) dataset focused on medicinal plants and herbs. It was developed for scientific research at Ton Duc Thang University (TDTU), with the goal of advancing AI models capable of recognizing and answering questions about Vietnamese medicinal herbs.
Homepage: Hugging Face Dataset
Repository: azan100an/tdtu_vqa_dataset_herb
Point of Contact:… See the full description on the dataset page: https://huggingface.co/datasets/azan100an/tdtu_vqa_dataset_herb.beaker-vqaist-vqa-zip-full
IST-VQA — Multilingual Scene Text VQA Dataset
Curated dataset for Indic Scene Text Visual Question Answering covering three languages:
Bengali (bn) · Hindi (hi) · Tamil (ta)
Summary
Bengali
Hindi
Tamil
Total
Total Images
2,308
2,097
2,322
6,727
Total VQA Pairs
2,306
2,097
2,321
6,724
— Crawl images
1,615
1,570
1,394
4,579
— IndicSTR12 images
357
173
336
866
— Synthetic images
336
354
592
1,282
Sources: Real-world web-crawled scene photos… See the full description on the dataset page: https://huggingface.co/datasets/somusan/ist-vqa-zip-full.critique-VQA-SFTfood-VQA-datasetVQA
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/sajjadrauf/VQA.vqav2-small-dataMixedWM38-VQA
Wafer VQA Dataset
Overview
Wafer VQA Dataset is a multimodal benchmark for wafer map understanding, visual question answering, and defect reasoning. It is built on top of the MixedWM38 wafer-map collection and organized into two annotation styles:
tuple_generation: one multi-question response per image, intended for GRPO or other sequence-level optimization settings
stepwise_reasoning: one stepwise dialogue per image, intended for supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/wafervqaanon/MixedWM38-VQA.sleap-mice-vqavqa-food-vietnamesevietnamese-medicinal-herb-vqa
Vietnamese Medicinal Herb VQA
This dataset contains Vietnamese visual question answering samples for medicinal herbs and traditional medicinal materials.
Each JSONL row has the following fields:
{
"id": "P000002676",
"name": "Cây Hẹ",
"image": "images/P000002266.jpg",
"question": "Hoa trong ảnh có màu gì?",
"answer": "Màu trắng"
}
Splits
Split
Images
QA pairs
Train
3032
23211
Validation
379
2998
Test
380
2952
The test split includes a… See the full description on the dataset page: https://huggingface.co/datasets/gdllelf/vietnamese-medicinal-herb-vqa.Vietnamese-Food-VQA
Vietnamese Food VQA Dataset
Bộ dữ liệu Hỏi đáp Hình ảnh (VQA) chuyên biệt cho các món ăn Việt Nam. Được xây dựng phục vụ cho mục đích nghiên cứu và học tập trong lĩnh vực Deep Learning.
Thông tin bộ dữ liệu
Ngôn ngữ: Tiếng Việt
Lĩnh vực: Ẩm thực Việt Nam (Phở, Bún bò, Bánh mì, Cơm tấm, v.v.)
Số lượng mẫu: ~2500 - 10000 QA pairs (bao gồm dữ liệu đã augmented)
Định dạng:
Ảnh: .jpg
Nhãn: .json (train, val, test)
Cấu trúc thư mục trên Hugging Face
data/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Hoppaaa/Vietnamese-Food-VQA.260514_fan_vqa_pre_proc
260514 Fan VQA Pre-Proc Dataset
VQA.
minecraft-vqaurdu-vqa-dataset
