datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3D-ADAMRepository for the 3D-ADAM (3D Anomaly Detection in Additive Manufacturing) Dataset. This is the raw data for our complete dataset, separated by part-instance to allow users to utilise the dataset as desired.
We provide a single-camera (using the MechMind-Nano) subset prepared for unsupervised training at anomaly detection, localisation and segementation tasks through the anomalib library in a separate repository: here
Our ArXiv paper can also be found here: 3D-ADAM Dataset
This project has… See the full description on the dataset page: https://huggingface.co/datasets/pmchard/3D-ADAM.PMC-VQA
PMC-VQA Dataset
PMC-VQA Dataset
Daraset Structure
Sample
Dataset Structure
PMC-VQA (version-1: 227k VQA pairs of 149k images).
train.csv: metafile of train set
test.csv: metafile of test set
test_clean.csv: metafile of test clean set
images.zip: images folder
(update version-2: noncompound images).
train2.csv: metafile of train set
test2.csv: metafile of test set
images2.zip: images folder
Sample
A row in train.csv is shown bellow… See the full description on the dataset page: https://huggingface.co/datasets/RadGenome/PMC-VQA.open-pmc-18m
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.PMC-Clinical-VQA
PMC-VQA: A Large-Scale Visual Question Answering Dataset for Clinical Figures
This dataset contains over 1,700,000 Visual Question Answering (VQA) samples derived from figures and charts in biomedical articles from PubMed Central (PMC).
This is a preliminary release. A full dataset card and an accompanying research paper are currently in preparation.
Raw version of this dataset with licenses and metadata can be found on Hugging Face: DermaVLM/pmc_clinical_VQA_raw
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DermaVLM/PMC-Clinical-VQA.PMC-VQA-1
PMC-VQA-1
This dataset is a streaming-friendly version of the PMC-VQA dataset, specifically containing the "Compounded Images" version (version-1). It is designed to facilitate efficient training and evaluation of Visual Question Answering (VQA) models in the medical domain, straight from the repository
Dataset Description
The original PMC-VQA dataset, available at https://huggingface.co/datasets/xmcmic/PMC-VQA, comprises Visual Question Answering pairs derived from… See the full description on the dataset page: https://huggingface.co/datasets/hamzamooraj99/PMC-VQA-1.PMC-VQA-refinedMedVLThinker-pmc_vqaCode: https://github.com/UCSC-VLAA/MedVLThinker
Project Page: https://ucsc-vlaa.github.io/MedVLThinker/
📊 Datasets
Available Datasets
Our project provides several curated datasets for medical vision-language understanding and training:
Dataset
Modality
Description
Download
MedVLThinker-m23k-tokenized
Text-only
Tokenized version of the m23k dataset
🤗 HF
MedVLThinker-pmc_vqa-gpt_4o_reasoning-tokenized
Image-Text
Tokenized PMC-VQA dataset with GPT-4o generated… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedVLThinker-pmc_vqa.open-pmc
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc.3D-ADAM_anomalibRepository for the 3D-ADAM (3D Anomaly Detection in Additive Manufacturing) Dataset.
This is a single-camera (using the MechMind-Nano) subset prepared for unsupervised training at anomaly detection, localisation and segementation tasks through the anomalib library.
The full training, test and validation sets can be viewed and interacted with in this repository, and we also provide a zipped copy of the data which can be accessed automatically when using the dataset through the anomalib… See the full description on the dataset page: https://huggingface.co/datasets/pmchard/3D-ADAM_anomalib.PMC-VQA_smallPMC-VQA
PMC-VQA - PubMed Central Visual Question Answering
Description
This dataset contains visual question answering data from PubMed Central medical literature. Questions are designed as multiple choice format requiring understanding of medical figures and images. 16 reasoning traces were collected for each example in this task by sampling with GPT-4o, available in the responses column. We greatly appreciate and build from the original data source available at… See the full description on the dataset page: https://huggingface.co/datasets/OctoMed/PMC-VQA.MedVLThinker-pmc_vqa-gpt_4o_reasoning-tokenizedCode: https://github.com/UCSC-VLAA/MedVLThinker
Project Page: https://ucsc-vlaa.github.io/MedVLThinker/
📊 Datasets
Available Datasets
Our project provides several curated datasets for medical vision-language understanding and training:
Dataset
Modality
Description
Download
MedVLThinker-m23k-tokenized
Text-only
Tokenized version of the m23k dataset
🤗 HF
MedVLThinker-pmc_vqa-gpt_4o_reasoning-tokenized
Image-Text
Tokenized PMC-VQA dataset with GPT-4o generated… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedVLThinker-pmc_vqa-gpt_4o_reasoning-tokenized.PMC-VQA-text
PMC-VQA-text
This dataset is a text format of PMC-VQA.
We built this dataset using the Meta-Llama-3-70B-Instruct, and the instruction we used is: Rewrite the question-answer pairs into a paragraph format (Do not use the words 'question' and 'answer' in your responses):.
train_text.json corresponds to the train.csv and train_2.csv splits in the PMC-VQA dataset.
Samples with two or more question-and-answer pairs were selected.
Citation
If you find this dataset useful… See the full description on the dataset page: https://huggingface.co/datasets/myeongkyunkang/PMC-VQA-text.pmc-vqa-v2pmc-vqa-robustnesspmc_derma_VQA_with_metadata_processedPMC-VQAbiomed-pmc-vqapmc-vqapmc_clinical_VQA_rawFormatted version of this dataset can be found on Hugging Face: DermaVLM/PMC-Clinical-VQA
PMC-VQA-2Fork of RadGenome/PMC-VQA converted to:
Use v2 of the PMC-VQA (train_2.csv and test_2.csv).
Wrap images as binary object
Remove spacing and option prefix (A:) from questions and choices
Metadata
Name
#train
#val
#test
img#train
img#val
img#test
PMC-VQA
152,603
0
33,430
135,339
0
29,021
Conversion script
from pathlib import Path
from datasets import Dataset, Features, Image, Value
SLAKE_FEAT = {
"index": Value("uint32"),
"image":… See the full description on the dataset page: https://huggingface.co/datasets/CAIR-M3LLM/PMC-VQA-2.train2-PMC_Derma_VQA
Asset from the SCALEMED Framework
This model/dataset is an asset released as part of the SCALEMED framework, a project focused on developing scalable and resource-efficient medical AI assistants.
Project Overview
The models, known as DermatoLlama, were trained on versions of the DermaSynth dataset, which was also generated using the SCALEMED pipeline.
For a complete overview of the project, including all related models, datasets, and the source code, please visit our main… See the full description on the dataset page: https://huggingface.co/datasets/DermaVLM/train2-PMC_Derma_VQA.pmc_oa_demoFoundation models trained on large-scale dataset gain a recent surge in CV and NLP. In contrast, development in biomedical domain lags far behind due to data scarcity.
To address this issue, we build and release PMC-OA, a biomedical dataset with 1.6M image-caption pairs collected from PubMedCentral's OpenAccess subset, which is 8 times larger than before.
PMC-OA covers diverse modalities or diseases, with majority of the image-caption samples aligned at finer-grained level, i.e., subfigure and subcaption.
While pretraining a CLIP-style model on PMC-OA, our model named PMC-CLIP achieves state-of-the-art results on various downstream tasks,
including image-text retrieval on ROCO, MedMNIST image classification, Medical VQA, i.e. +8.1% R@10 on image-text retrieval, +3.9% accuracy on image classification.open-pmc-18m
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/chjudy/open-pmc-18m.PMC-VQARL_Mixed_MCQ_Filtered_Our_PMCRL_Mixed_MCQ_Filtered_Our_PMC_New_Image_Tagpmc_derma_VQA_with_metadataopen-pmc-18m-subsetpmc_vqa_test
