datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya-mm-exams-spanish-medicalMedical Spanish Exams for the Multimodal Aya Exams Projects.
Questions available in file: data.json
Images stored in: /images
Original data and file available here: link
GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.GMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.Japanese-Medical-VQA-12m
Japanese Medical VQA 12M
Japanese Medical VQA 12M is a large-scale Japanese medical multimodal dataset built from Open-PMC-18M and released in Parquet and Webdataset format.
This dataset contains outputs from multiple data-construction stages, including:
source captions
Japanese translations of source captions
enriched captions
Japanese translations of enriched captions
question-answering
Current Repository Format
This repository currently stores the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/MIL-UT/Japanese-Medical-VQA-12m.synthetic-medical-document-recognition-benchmark
Synthetic Medical Document Recognition Benchmark
This dataset contains synthetic, English-language medical records rendered as
documents for evaluating automated data extraction and de-identification
systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple
visual representations derived from that record.
Every rendered document is clearly marked as synthetic. This makes the dataset
suitable for manual testing, product demonstrations, and workflows that… See the full description on the dataset page: https://huggingface.co/datasets/morzel85/synthetic-medical-document-recognition-benchmark.medical-prescription-datasetmimic-medical-imaging-qa
MIMIC Medical Imaging QA Dataset
5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction.
License
The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.medical-prescription-dataset100-handwritten-medical-recordsmedical-vqamotor-medicina-legala-v2-storagemedical-vqa-8-datasetsProject-Imaging-X
Project Imaging-X is a strategic initiative to consolidate 1000+ open medical imaging datasets worldwide, breaking down data silos through systematic integration to build the foundational infrastructure for next-generation medical AI models.
Challenge: Medical imaging lacks large-scale unified datasets due to clinical expertise requirements and privacy constraints, limiting the development of powerful medical foundation models.
Solution: We surveyed 1000+… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/Project-Imaging-X.MEDIC
MEDIC: A Multi-Task Learning Dataset for Disaster Image Classification
Data
The MEDIC is the largest multi-task learning disaster-related dataset, an extended version of the crisis image benchmark dataset. It consists of data from several sources, including CrisisMMD, data from AIDR, and the Damage Multimodal Dataset (DMD). The dataset contains 71,198 images.
Data Format and Directories
Directories
data: Main directory with the following… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MEDIC.medical_recipe_datasetThis is fully synthetic! Это полная синтетика!
Форма № 107-1/у
Medical-OCR-Training
Handwritten OCR Dataset
Total files - ~27k. Split into folders to comply with HF norms (10k limit per folder, 25k limit per commit)
Need to update paths in JSON and CSV files
Medical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.MedicalHistoryofBritishIndia
A Medical History of British India Dataset
Dataset Description
This dataset contains digitized official publications documenting medical research and public health in British India from 1850-1950. The collection represents a crucial period in medical history, capturing the transition from humoral to biochemical medical traditions and documenting major breakthroughs in bacteriology, parasitology, and vaccine development. These documents provide invaluable insights into… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/MedicalHistoryofBritishIndia.noisy-medical-document-images-ocr
🏥 Noisy Medical Document Images (OCR)
1,000 noisy, synthetic medical document images with structured JSON ground truth — built for Document AI, LayoutLM fine-tuning, and clinical NLP research.
🧭 Overview
This dataset provides 1,000 high-resolution images of two healthcare document types, each degraded with realistic scanning artifacts to simulate real-world OCR conditions:
Category
Count
Description
🧾 Hospital Bills
500
Itemized statements with… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/noisy-medical-document-images-ocr.chest_no_bb_with_meta_v2medical-rl-grpo-v1medical_anomalies_8medical-prescriptions_beirThis is a copy of https://huggingface.co/datasets/jinaai/medical-prescriptions reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/medical-prescriptions_beir.medical-history-of-british-india
A Medical History of British India Dataset
Dataset Description
This dataset contains digitiaed official publications documenting medical research and public health in British India from 1850-1950. The collection represents a crucial period in medical history, capturing the transition from humoral to biochemical medical traditions and documenting major breakthroughs in bacteriology, parasitology, and vaccine development. These documents provide invaluable insights into… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/medical-history-of-british-india.Italian-OTC-medicines
Italian-OTC-medicines - Over-the-Counter Drug Packaging Dataset
Overview
This dataset consists of high-resolution images of Over-the-Counter (OTC) medication packaging primarily from the Italian market. It is designed for researchers and developers working on computer vision tasks such as Optical Character Recognition (OCR), Object Detection, and Information Extraction within the pharmaceutical and healthcare sectors.
Each image in the dataset is paired with a… See the full description on the dataset page: https://huggingface.co/datasets/mennox/Italian-OTC-medicines.OpenMM_Medical
OpenMM-Medical
Introduction
OpenMM-Medical is a comprehensive medical evaluation dataset, which is an integration of existing datasets. OpenMM-Medical spans multiple domains, including Magnetic Resonance Imaging (MRI), CT scans, X-rays, microscopy images, endoscopy, fundus imaging, and dermoscopy.
Components
Content
Type
Number
Metrics
ACRIMA
Fundus Photography
Multiple Choice Question Answering
159
Acc
Adam Challenge
Endoscopy
Multiple Choice Question… See the full description on the dataset page: https://huggingface.co/datasets/baichuan-inc/OpenMM_Medical.medical_anomalies_0-3Medical-ROIs-K2.6
Medical-ROIs-K2.6
Medical visual grounding SFT data: for each clinical VQA sample, a teacher model proposes answer-supporting ROIs (regions of interest) as 2D bounding boxes.
Teacher: moonshotai/Kimi-K2.6.Upstream images & QA: MBZUAI/medix-rl-data.
How the data is generated
MBZUAI/medix-rl-data (train)
│
│ problem / solution / image / source / id
▼
Teacher: moonshotai/Kimi-K2.6
(vision + text; given question + gold answer)
│… See the full description on the dataset page: https://huggingface.co/datasets/erow/Medical-ROIs-K2.6.medicinal_plant_classification_bd
Medicinal Plant Classification Bd
This dataset contains real RGB images of medicinal plants native to Bangladesh, captured in a controlled laboratory environment. Images were collected using handheld smartphones during the summer months (July to August), providing a diverse and standardized representation of plant specimens under consistent lighting and background conditions. The dataset contains 5,000 images across 10 classes: Bohera, Devilbackbone, Haritoki, Lemongrass… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/medicinal_plant_classification_bd.
