CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lmms-lab-encoder /VQAv2image100K<n<1M38 likes36k downloads3y agoHugging Face02ChongyanChen /VQAonline VQAonline 🌐 Homepage | 🤗 Dataset | 📖 arXiv Dataset Description We introduce VQAonline, the first VQA dataset in which all contents originate from an authentic use case. VQAonline includes 64K visual questions sourced from an online question answering community (i.e., StackExchange). It differs from prior datasets; examples include that it contains: (1) authentic context that clarifies the question (2) an answer the individual asking the question validated as… See the full description on the dataset page: https://huggingface.co/datasets/ChongyanChen/VQAonline.imagevisual-question-answering10K<n<100K16 likes20k downloads2y agoHugging Face03Dangindev /viet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage. It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English. The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing, landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life, and transportation.visual-question-answering10K<n<100K8 likes14k downloads10mo agoHugging Face04SimulaMet /Kvasir-VQA-x1 Kvasir-VQA-x1 A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy Kvasir-VQA-x1 on GitHub | Original Image from Kvasir-VQA(Simula Datasets) | Paper 🔗 MediaEval Medico 2025 Challenge uses this dataset. We encourage you to check out and participate! Overview Kvasir-VQA-x1 is a large-scale dataset designed to benchmark medical visual question answering (MedVQA) in gastrointestinal (GI) endoscopy. It introduces 159,549 new QA… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/Kvasir-VQA-x1.imagevisual-question-answering100K<n<1M16 likes10k downloads1y agoHugging Face05VQA-Illusion /IllusionChar_train IllusionChar — Training Set Dataset summary This repository contains the training split of IllusionChar, the optical character recognition (OCR) component of Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. The task is to transcribe a hidden, case-sensitive alphanumeric sequence from an illusory image, or return No illusion when no sequence is embedded. Sequences contain 3–5 characters drawn from digits, uppercase Latin letters, and… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionChar_train.imageimage-to-text10K<n<100K1 likes8.9k downloads19d agoHugging Face06lmms-lab-encoder /VizWiz-VQA Dataset Card for "VizWiz-VQA" Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of VizWiz-VQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{gurari2018vizwiz, title={Vizwiz grand… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/VizWiz-VQA.image10K<n<100K9 likes8.7k downloads3y agoHugging Face07flaviagiammarino /vqa-rad Dataset Card for VQA-RAD Dataset Description VQA-RAD is a dataset of question-answer pairs on radiology images. The dataset is intended to be used for training and testing Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions. The dataset is built from MedPix, which is a free open-access online database of medical images. The question-answer pairs were manually generated by a team of clinicians.… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/vqa-rad.imagevisual-question-answering1K<n<10K102 likes7.2k downloads3y agoHugging Face08VQA-Illusion /IllusionChar_test IllusionChar — Test Set Dataset summary This repository contains the public test split of IllusionChar, the OCR benchmark introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each metadata row can be paired across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control conditions. The expected output for an illusion-bearing or source-condition image is an exact, case-sensitive… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionChar_test.imageimage-to-text10K<n<100K0 likes6.9k downloads19d agoHugging Face09flaviagiammarino /path-vqa Dataset Card for PathVQA Dataset Description PathVQA is a dataset of question-answer pairs on pathology images. The dataset is intended to be used for training and testing Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions. The dataset is built from two publicly-available pathology textbooks: "Textbook of Pathology" and "Basic Pathology", and a publicly-available digital library: "Pathology… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/path-vqa.imagevisual-question-answering10K<n<100K75 likes5.7k downloads3y agoHugging Face10lmms-lab-encoder /OK-VQAimage1K<n<10K8 likes5.5k downloads3y agoHugging Face11VQA-Illusion /MNIST_train IllusionMNIST — Training Set Dataset summary This repository contains the training split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. The dataset is intended for training models to recognize MNIST digits embedded as visual illusions (pareidolia) in generated scenes and to reject images that contain no illusion. MNIST source-condition images were sampled and resized to 512 × 512 pixels, combined… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_train.imageimage-classification1K<n<10K0 likes5k downloads19d agoHugging Face12VQA-Illusion /FashionMnist_train IllusionFashionMNIST — Training Set Dataset summary This repository contains the training split of IllusionFashionMNIST, one of the four datasets introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. It is designed to train and evaluate models on the recognition of Fashion-MNIST categories embedded as visual illusions (pareidolia) in generated scenes. The source-condition images are sampled from Fashion-MNIST and resized to… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/FashionMnist_train.imageimage-classification1K<n<10K1 likes4.7k downloads19d agoHugging Face13inclusionAI /ZwZ-RL-VQA ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception This synthetic dataset is generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use. 📖 Overview The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive: Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V)… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZwZ-RL-VQA.text100K<n<1M17 likes4.5k downloads4mo agoHugging Face14VQA-Illusion /IllusionAnimals_train IllusionAnimals — Training Set Dataset summary This repository contains the training split of IllusionAnimals, one of the four benchmarks introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. It supports training models to identify animal categories embedded as visual illusions (pareidolia) in generated scenes and to recognize when no illusion is present. The source-condition animal images were generated with SDXL-Lightning.… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionAnimals_train.imageimage-classification1K<n<10K1 likes4.5k downloads19d agoHugging Face15howard-hou /OCR-VQA Dataset Card for "OCR-VQA" More Information needed image100K<n<1M60 likes4.2k downloads3y agoHugging Face16VQA-Illusion /FashionMnist_test IllusionFashionMNIST — Test Set Dataset summary This repository contains the public test split of IllusionFashionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each metadata row identifies a Fashion-MNIST target and can be paired across five image conditions: source-condition, illusion, filtered illusion, illusionless control, and filtered illusionless control. The source-condition images originate from… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/FashionMnist_test.imageimage-classification1K<n<10K0 likes4.2k downloads19d agoHugging Face17LLDDSS /Awesome_Spatial_VQA_Benchmarksimage10K<n<100K1 likes4.1k downloads1y agoHugging Face18VQA-Illusion /MNIST_test IllusionMNIST — Test Set Dataset summary This repository contains the public test split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Every indexed example can be compared across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control images. The source-condition images are sampled from MNIST and resized to 512 × 512 pixels. Illusion images were… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_test.imageimage-classification1K<n<10K0 likes4k downloads19d agoHugging Face19VQA-Illusion /IllusionAnimals_test IllusionAnimals — Test Set Dataset summary This repository contains the public test split of IllusionAnimals, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each annotated example is paired across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control conditions. The animal source-condition images were generated with SDXL-Lightning. English scene descriptions and… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionAnimals_test.imageimage-classification1K<n<10K2 likes3.8k downloads19d agoHugging Face20tumor-vqa /DeepTumorVQA_2.0 DeepTumorVQA v2 3D abdominal-CT diagnostic Visual Question Answering benchmark with 42 clinical subtypes and 438K total QA pairs (10K curated benchmark + 428K training pool). Includes pre-extracted 2D and video modalities, 20K agent training trajectories with tool-use traces, and a paper-locked leaderboard. Resources 📄 Paper (arXiv) https://arxiv.org/abs/2605.09679 💻 Code (GitHub) https://github.com/Schuture/DeepTumorVQA 🤗 Dataset (this… See the full description on the dataset page: https://huggingface.co/datasets/tumor-vqa/DeepTumorVQA_2.0.imagevisual-question-answering10K<n<100K6 likes3.4k downloads4mo agoHugging Face21MIL-UT /Japanese-Medical-VQA-12m Japanese Medical VQA 12M Japanese Medical VQA 12M is a large-scale Japanese medical multimodal dataset built from Open-PMC-18M and released in Parquet and Webdataset format. This dataset contains outputs from multiple data-construction stages, including: source captions Japanese translations of source captions enriched captions Japanese translations of enriched captions question-answering Current Repository Format This repository currently stores the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/MIL-UT/Japanese-Medical-VQA-12m.imageimage-to-text10M<n<100M7 likes3.3k downloads6mo agoHugging Face22OSU-AIoT-MLSys-Lab /SuperMemory-VQA SuperMemoryVQA SuperMemory-VQA is an egocentric visual question answering benchmark for evaluating long-horizon memory in augmented reality assistant settings. The dataset is designed around practical questions a person might ask a wearable memory assistant, such as where an object was left, what someone said earlier, whether a planned step was completed, or what happened next in a longer event. The benchmark contains 4,853 human-verified question-answer pairs grounded in 52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.tabularvisual-question-answering1K<n<10K5 likes3.2k downloads3mo agoHugging Face23merve /vqav2-smallimage10K<n<100K21 likes2.9k downloads2y agoHugging Face24RadGenome /PMC-VQA PMC-VQA Dataset PMC-VQA Dataset Daraset Structure Sample Dataset Structure PMC-VQA (version-1: 227k VQA pairs of 149k images). train.csv: metafile of train set test.csv: metafile of test set test_clean.csv: metafile of test clean set images.zip: images folder (update version-2: noncompound images). train2.csv: metafile of train set test2.csv: metafile of test set images2.zip: images folder Sample A row in train.csv is shown bellow… See the full description on the dataset page: https://huggingface.co/datasets/RadGenome/PMC-VQA.image79 likes2.8k downloads2y agoHugging Face25mmoukouba /MedPix-VQA MedPix-VQA Dataset The MedPix-VQA dataset is a version of the data found at MEDPIX-ClinQA, specifically modified to address an image overlap issue that would result from directl splitting the original dataset. This overlap can lead to a model potentially seeing the same image during both training and validation, potentially leading to bias or data leakage. Key Modifications: We have modified the dataset to ensure no image overlap between the training and validation… See the full description on the dataset page: https://huggingface.co/datasets/mmoukouba/MedPix-VQA.image10K<n<100K0 likes2.6k downloads1y agoHugging Face26SII-Monument-Valley /CiQi-VQA CiQi-Agent Github | Model | Dataset | Paper CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains Accepted to ECCV 2026 🎯 Overview CiQi-Agent has been accepted to ECCV 2026. We present CiQi-Agent, a domain-specific multimodal agent for antique Chinese porcelain connoisseurship. The project is designed to combine fine-grained visual perception, tool-augmented reasoning, and cultural-heritage knowledge… See the full description on the dataset page: https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA.imagequestion-answering10K<n<100K6 likes2k downloads28d agoHugging Face27orena-dkfz /heico-focus-vqagated HeiCo-FOCUS (Beta release) A clinically grounded dataset for long-context video understanding in minimally invasive surgery. 📄 Paper &nbsp;•&nbsp; 🤗 Dataset &nbsp;•&nbsp; 💻 Code &nbsp;•&nbsp; 🏆 Challenge &nbsp;•&nbsp; ⚖️ CC BY-NC-SA 4.0 [!NOTE] Until 31 July 2026, we will employ a restricted post-release review period. During this time, we kindly ask the community to provide feedback to help us increase quality control and make continuous improvements.… See the full description on the dataset page: https://huggingface.co/datasets/orena-dkfz/heico-focus-vqa.textvisual-question-answering10K<n<100K12 likes2k downloads2mo agoHugging Face28mlinhbng /viet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage. It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English. The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing, landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life, and transportation.visual-question-answering10K<n<100K0 likes1.9k downloads10mo agoHugging Face29worldcuisines /vqa WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines This version includes all images in the dataset. For a more lightweight and accessible alternative, please refer to the (1.1 release)[https://huggingface.co/datasets/worldcuisines/vqa-v1.1/] which reduces download size while preserving all text and metadata. The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆. WorldCuisines is a… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa.image1M<n<10M30 likes1.8k downloads10mo agoHugging Face30Multimodal-Fatima /VQAv2_train Dataset Card for "VQAv2_train" More Information needed image100K<n<1M5 likes1.8k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.