datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.medpmc-screening-dataset
MedPMC Initial Screening Training/Test Datasets
Overview
This dataset contains the annotations used for the initial screening stage of the MedPMC framework, which aims to identify clinically relevant medical images from biomedical literature.
The training and validation sets are automatically curated using GPT-4o. The test set is manually annotated.
For details on dataset construction, annotation guidelines, and the overall MedPMC pipeline, please refer to our… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-screening-dataset.medpmc-multi-fig-detection-dataset
MedPMC Multipanel Figure Detection Training/Test Datasets
Overview
This dataset contains the annotations used for the multipanel figure detection stage of the MedPMC framework. The task is to identify whether a biomedical figure is a multipanel, or compound, figure. This stage supports downstream subfigure extraction and caption alignment by separating figures that contain multiple visual panels from standalone figures.
The training and validation sets combine… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-multi-fig-detection-dataset.M3LLM-data-release
