CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lmms-lab-encoder /VQAv2image100K<n<1M38 likes35k downloads3y agoHugging Face02ChongyanChen /VQAonline VQAonline 🌐 Homepage | 🤗 Dataset | 📖 arXiv Dataset Description We introduce VQAonline, the first VQA dataset in which all contents originate from an authentic use case. VQAonline includes 64K visual questions sourced from an online question answering community (i.e., StackExchange). It differs from prior datasets; examples include that it contains: (1) authentic context that clarifies the question (2) an answer the individual asking the question validated as… See the full description on the dataset page: https://huggingface.co/datasets/ChongyanChen/VQAonline.imagevisual-question-answering10K<n<100K16 likes21k downloads2y agoHugging Face03SimulaMet /Kvasir-VQA-x1 Kvasir-VQA-x1 A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy Kvasir-VQA-x1 on GitHub | Original Image from Kvasir-VQA(Simula Datasets) | Paper 🔗 MediaEval Medico 2025 Challenge uses this dataset. We encourage you to check out and participate! Overview Kvasir-VQA-x1 is a large-scale dataset designed to benchmark medical visual question answering (MedVQA) in gastrointestinal (GI) endoscopy. It introduces 159,549 new QA… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/Kvasir-VQA-x1.imagevisual-question-answering100K<n<1M16 likes10k downloads1y agoHugging Face04lmms-lab-encoder /VizWiz-VQA Dataset Card for "VizWiz-VQA" Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of VizWiz-VQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{gurari2018vizwiz, title={Vizwiz grand… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/VizWiz-VQA.image10K<n<100K9 likes8.7k downloads3y agoHugging Face05flaviagiammarino /vqa-rad Dataset Card for VQA-RAD Dataset Description VQA-RAD is a dataset of question-answer pairs on radiology images. The dataset is intended to be used for training and testing Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions. The dataset is built from MedPix, which is a free open-access online database of medical images. The question-answer pairs were manually generated by a team of clinicians.… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/vqa-rad.imagevisual-question-answering1K<n<10K104 likes7.3k downloads3y agoHugging Face06flaviagiammarino /path-vqa Dataset Card for PathVQA Dataset Description PathVQA is a dataset of question-answer pairs on pathology images. The dataset is intended to be used for training and testing Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions. The dataset is built from two publicly-available pathology textbooks: "Textbook of Pathology" and "Basic Pathology", and a publicly-available digital library: "Pathology… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/path-vqa.imagevisual-question-answering10K<n<100K75 likes5.9k downloads3y agoHugging Face07lmms-lab-encoder /OK-VQAimage1K<n<10K8 likes5.3k downloads3y agoHugging Face08LLDDSS /Awesome_Spatial_VQA_Benchmarksimage10K<n<100K1 likes4.5k downloads1y agoHugging Face09inclusionAI /ZwZ-RL-VQA ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception This synthetic dataset is generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use. 📖 Overview The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive: Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V)… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZwZ-RL-VQA.text100K<n<1M17 likes4.3k downloads4mo agoHugging Face10howard-hou /OCR-VQA Dataset Card for "OCR-VQA" More Information needed image100K<n<1M60 likes4k downloads3y agoHugging Face11OSU-AIoT-MLSys-Lab /SuperMemory-VQA SuperMemoryVQA SuperMemory-VQA is an egocentric visual question answering benchmark for evaluating long-horizon memory in augmented reality assistant settings. The dataset is designed around practical questions a person might ask a wearable memory assistant, such as where an object was left, what someone said earlier, whether a planned step was completed, or what happened next in a longer event. The benchmark contains 4,853 human-verified question-answer pairs grounded in 52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.tabularvisual-question-answering1K<n<10K5 likes3.4k downloads3mo agoHugging Face12tumor-vqa /DeepTumorVQA_2.0 DeepTumorVQA v2 3D abdominal-CT diagnostic Visual Question Answering benchmark with 42 clinical subtypes and 438K total QA pairs (10K curated benchmark + 428K training pool). Includes pre-extracted 2D and video modalities, 20K agent training trajectories with tool-use traces, and a paper-locked leaderboard. Resources 📄 Paper (arXiv) https://arxiv.org/abs/2605.09679 💻 Code (GitHub) https://github.com/Schuture/DeepTumorVQA 🤗 Dataset (this… See the full description on the dataset page: https://huggingface.co/datasets/tumor-vqa/DeepTumorVQA_2.0.imagevisual-question-answering10K<n<100K6 likes3.3k downloads4mo agoHugging Face13merve /vqav2-smallimage10K<n<100K21 likes3k downloads2y agoHugging Face14MIL-UT /Japanese-Medical-VQA-12m Japanese Medical VQA 12M Japanese Medical VQA 12M is a large-scale Japanese medical multimodal dataset built from Open-PMC-18M and released in Parquet and Webdataset format. This dataset contains outputs from multiple data-construction stages, including: source captions Japanese translations of source captions enriched captions Japanese translations of enriched captions question-answering Current Repository Format This repository currently stores the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/MIL-UT/Japanese-Medical-VQA-12m.imageimage-to-text10M<n<100M7 likes2.8k downloads6mo agoHugging Face15Multimodal-Fatima /VQAv2_train Dataset Card for "VQAv2_train" More Information needed image100K<n<1M5 likes1.8k downloads3y agoHugging Face16mmoukouba /MedPix-VQA MedPix-VQA Dataset The MedPix-VQA dataset is a version of the data found at MEDPIX-ClinQA, specifically modified to address an image overlap issue that would result from directl splitting the original dataset. This overlap can lead to a model potentially seeing the same image during both training and validation, potentially leading to bias or data leakage. Key Modifications: We have modified the dataset to ensure no image overlap between the training and validation… See the full description on the dataset page: https://huggingface.co/datasets/mmoukouba/MedPix-VQA.image10K<n<100K0 likes1.8k downloads1y agoHugging Face17worldcuisines /vqa WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines This version includes all images in the dataset. For a more lightweight and accessible alternative, please refer to the (1.1 release)[https://huggingface.co/datasets/worldcuisines/vqa-v1.1/] which reduces download size while preserving all text and metadata. The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆. WorldCuisines is a… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa.image1M<n<10M30 likes1.7k downloads10mo agoHugging Face18orena-dkfz /heico-focus-vqagated HeiCo-FOCUS (Beta release) A clinically grounded dataset for long-context video understanding in minimally invasive surgery. 📄 Paper &nbsp;•&nbsp; 🤗 Dataset &nbsp;•&nbsp; 💻 Code &nbsp;•&nbsp; 🏆 Challenge &nbsp;•&nbsp; ⚖️ CC BY-NC-SA 4.0 [!NOTE] Until 31 July 2026, we will employ a restricted post-release review period. During this time, we kindly ask the community to provide feedback to help us increase quality control and make continuous improvements.… See the full description on the dataset page: https://huggingface.co/datasets/orena-dkfz/heico-focus-vqa.textvisual-question-answering10K<n<100K12 likes1.7k downloads2mo agoHugging Face19mdwiratathya /SLAKE-vqa-english Dataset Description SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering [ISBI 2021 oral] Corresponding Authors: Bo Liu, Xiao-Ming Wu Original dataset is retrieved from https://huggingface.co/datasets/BoKelvin/SLAKE In this dataset, we modified some things to match our task: The original dataset are bilingual, we filtered to only English We only take the image (converted as PIL object), question, and answer column Any questions, please… See the full description on the dataset page: https://huggingface.co/datasets/mdwiratathya/SLAKE-vqa-english.image1K<n<10K5 likes1.5k downloads2y agoHugging Face20liyf001 /CT-RATE-VQA CT-RATE-VQA Dataset We constructed a large-scale CT-VQA dataset based on the ReXGroundingCT data \cite{rexct} to support model training and evaluation.For each case, CT volumes were processed along with their corresponding multi-class segmentation masks, where each mask channel represents a specific lesion type. This dataset is designed for medical visual question answering (Med-VQA) tasks. image10K<n<100K1 likes1.4k downloads1y agoHugging Face21DermaVLM /PMC-Clinical-VQA PMC-VQA: A Large-Scale Visual Question Answering Dataset for Clinical Figures This dataset contains over 1,700,000 Visual Question Answering (VQA) samples derived from figures and charts in biomedical articles from PubMed Central (PMC). This is a preliminary release. A full dataset card and an accompanying research paper are currently in preparation. Raw version of this dataset with licenses and metadata can be found on Hugging Face: DermaVLM/pmc_clinical_VQA_raw Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DermaVLM/PMC-Clinical-VQA.imagevisual-question-answering1M<n<10M1 likes1.3k downloads1y agoHugging Face22remyxai /vqasynth_spacellava VQASynth_spacellava Uses the VQASynth pipeline to synthesize spatialVQA samples, mixed with general VQA samples used to fine-tune LLaVA-v1.5-13b. imagevisual-question-answering10K<n<100K14 likes1.2k downloads2y agoHugging Face23orena-dkfz /lapchole-focus-vqagated LapChole-FOCUS-VQA A clinically grounded benchmark for long-context video understanding in minimally invasive surgery. 💻 Code &nbsp;•&nbsp; 🏆 Challenge &nbsp;•&nbsp; ⚖️ Data Usage Agreement [!IMPORTANT] 🔒 This is a gated dataset Access is granted only to participants of the ORena FOCUS Challenge and is subject to manual review. To be approved you must: Have a Hugging Face account and be logged in — downloads are only enabled for registered, authenticated… See the full description on the dataset page: https://huggingface.co/datasets/orena-dkfz/lapchole-focus-vqa.textvisual-question-answering10K<n<100K6 likes1.2k downloads2mo agoHugging Face24ll-13 /LRS-VQA LRS-VQA Dataset This repository contains the LRS-VQA benchmark dataset, presented in the paper When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning. Code: The associated code and evaluation scripts can be found on the project's GitHub repository: https://github.com/VisionXLab/LRS-VQA Introduction Efficient vision-language understanding of large Remote Sensing Images (RSIs) is meaningful but challenging. Current… See the full description on the dataset page: https://huggingface.co/datasets/ll-13/LRS-VQA.textimage-text-to-text1K<n<10K2 likes1k downloads1y agoHugging Face25raidium /RadImageNet-VQAgated RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering We introduce RadImageNet-VQA, a large-scale dataset designed for training and benchmarking radiologic VQA on CT and MRI exams. Built from the CT/MRI subset of RadImageNet and its expert-curated anatomical and pathological annotations, RadImageNet-VQA provides 750K images with 7.5M generated samples, including 750K medical captions for visual-text alignment and 6.75M… See the full description on the dataset page: https://huggingface.co/datasets/raidium/RadImageNet-VQA.imagevisual-question-answering1M<n<10M92 likes960 downloads3mo agoHugging Face26siyrus /BToks-vidore_rag_Infographic-VQA BToks ViDoRe Infographic-VQA This dataset repository contains Lance-format converted data used by the open-source reproduction code for Bottleneck Tokens for Unified Multimodal Retrieval (arXiv:2604.11095). Source Converted from vidore/colpali_train_set. Subset/view: infographic_vqa. This repository does not change upstream ownership, licensing, citation requirements, or usage restrictions. Format The data is stored as Lance tables for the… See the full description on the dataset page: https://huggingface.co/datasets/siyrus/BToks-vidore_rag_Infographic-VQA.imageimage-to-text10K<n<100K0 likes953 downloads3mo agoHugging Face27Intel /SK-VQA Dataset Card for SQ-VQA Dataset Summary SK-VQA is a large-scale synthetic multimodal dataset containing over 2 million visual question-answer pairs, each paired with context documents that contain the information needed to answer the questions. The dataset is designed to address the critical need for training and evaluating multimodal LLMs (MLLMs) in context-augmented generation settings, particularly for retrieval-augmented generation (RAG) systems. It enables training… See the full description on the dataset page: https://huggingface.co/datasets/Intel/SK-VQA.image1M<n<10M1 likes948 downloads1y agoHugging Face28RL-MIND /HARD-VQA HARD-VQA Ultra-High-Resolution Aerial VQA with Original Images Embedded per Question 🤗 Hugging Face · 🟣 ModelScope · 📊 Statistics English | 中文:Hugging Face · ModelScope 📚 Introduction HARD-VQA packages ultra-high-resolution aerial visual question answering data as self-contained Parquet shards. Each row is one multiple-choice question, with all required original JPEG bytes embedded in its ordered images… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/HARD-VQA.imagevisual-question-answering1K<n<10K0 likes921 downloads6d agoHugging Face29hamzamooraj99 /PMC-VQA-1 PMC-VQA-1 This dataset is a streaming-friendly version of the PMC-VQA dataset, specifically containing the "Compounded Images" version (version-1). It is designed to facilitate efficient training and evaluation of Visual Question Answering (VQA) models in the medical domain, straight from the repository Dataset Description The original PMC-VQA dataset, available at https://huggingface.co/datasets/xmcmic/PMC-VQA, comprises Visual Question Answering pairs derived from… See the full description on the dataset page: https://huggingface.co/datasets/hamzamooraj99/PMC-VQA-1.imagevisual-question-answering100K<n<1M4 likes812 downloads2y agoHugging Face30LLDDSS /Awesome_Spatial_VQA_Benchmarks_ViewSpatial-Benchimage1K<n<10K0 likes798 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.