datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CVPR-BiomedSegFMThis repository contains the BiomedSegFM dataset, a crucial resource for the CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation.
Foundation Models for Interactive 3D Biomedical Image Segmentation (Homepage)
Foundation Models for Text-guided 3D Biomedical Image Segmentation (Homepage)
CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation
Highly recommend watching the webinar recording to learn about the task settings and… See the full description on the dataset page: https://huggingface.co/datasets/junma/CVPR-BiomedSegFM.biomedica_webdataset_24M
Dataset Card for Dataset Name
Arxiv: Arxiv
|
Website: Biomedica
|
Training instructions: OpenCLIP
|
Tutorial: Google Colab
BIOMEDICA Dataset is a large-scale, deep-learning-ready biomedical dataset containing over 24M imagecaption pairs and 30M image-references from 6M unique open-source articles. Each… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/biomedica_webdataset_24M.SPACCC_Tokenizer
The Tokenizer for Clinical Cases Written in Spanish
Introduction
This repository contains the tokenization model trained using the SPACCC_TOKEN corpus (https://github.com/PlanTL-SANIDAD/SPACCC_TOKEN). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to tokenize biomedical documents, specially clinical cases written in Spanish.
This model was created using the Apache… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Tokenizer.biomedical_lectures_v2
Vidore Benchmark 2 - MIT Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_v2.CodiEsp_corpus
Introduction
These are the train, development, test and background sets of the CodiEsp corpus. Train and development have gold standard annotations. The unannotated background and test sets are distributed together. All documents are released in the context of the CodiEsp track for CLEF ehealth 2020 (http://temu.bsc.es/codiesp/).
The CodiEsp corpus contains manually coded clinical cases. All documents are in Spanish language and CIE10 is the coding terminology (it is the Spanish… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/CodiEsp_corpus.Biomed-Enriched
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
Dataset Authors
Rian Touchent, Nathan Godey & Eric de la ClergerieSorbonne Université, INRIA Paris
Overview
Biomed-Enriched is a PubMed-derived dataset created using a two-stage annotation process. Initially, Llama 3.1 70B Instruct annotated 400K paragraphs for document type, domain, and educational quality. These annotations were then… See the full description on the dataset page: https://huggingface.co/datasets/almanach/Biomed-Enriched.BioMed-R1-Eval
Disentangling Reasoning and Knowledge in Medical Large Language Models
This is the evaluation dataset accompanying our paper, comprising 11 publicly available biomedical benchmarks. We disentangle each benchmark question into either medical reasoning or medical knowledge categories.
Additionally, we provide a set of adversarial reasoning traces designed to evaluate the robustness of medical reasoning models.
For more details, please refer to our GitHub.
If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/zou-lab/BioMed-R1-Eval.ProfNER_corpus_classificationProfNER_corpus_NER
Description
Gold standard annotations for profession detection in Spanish COVID-19 tweets
The entire corpus contains 10,000 annotated tweets. It has been split into training, validation, and test (60-20-20). The current version contains the training and development set of the shared task with Gold Standard annotations. In addition, it contains the unannotated test, and background sets will be released.
For Named Entity Recognition, profession detection, annotations are distributed… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/ProfNER_corpus_NER.Medical_Segmentation_Decathlon
🏆 Medical Segmentation Decathlon Dataset
📝 Overview
The Medical Segmentation Decathlon (MSD) is a comprehensive benchmark dataset for validating algorithms in 3D medical image segmentation. It includes 10 distinct tasks, each with unique challenges like small data sizes, unbalanced labels, varying object scales, multi-class labels, and multimodal imaging.
🔗 Dataset Access
🌐 Website: Medical Decathlon
📂 Google Drive: MSD Google Drive
🧩 Task… See the full description on the dataset page: https://huggingface.co/datasets/Novel-BioMedAI/Medical_Segmentation_Decathlon.SPACCC_Sentence-Splitter
The Sentence Splitter (SS) for Clinical Cases Written in Spanish
Introduction
This repository contains the sentence splitting model trained using the SPACCC_SPLIT corpus (https://github.com/PlanTL-SANIDAD/SPACCC_SPLIT). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to split sentences in biomedical documents, specially clinical cases written in Spanish. This model… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Sentence-Splitter.Biomedica2025EvalSet
Biomedica 2025 Eval Set
Unified test snapshot of the BioMedica 2025 vision–language evaluation
suites used in AMInZeroShotOpenEvalAllTasks. Every row is a single image
with closed-ended options, the gold answer, and provenance fields that
point back to the original dataset.
Images are stored as original JPEG/PNG bytes (or JPEG-encoded arrays) inside
parquet so the Hugging Face dataset viewer is enabled
(~1959 MB download, 89941 examples).
Suites
config… See the full description on the dataset page: https://huggingface.co/datasets/Alejandro98/Biomedica2025EvalSet.pubmed2024_sentence_embeddingsHLE-BioMedX
HLE-BioMedX — Multilingual HLE Biology/Medicine
A multilingual version of the Biology/Medicine subset of Humanity's Last Exam
(HLE), released as one subset per language.
Source benchmark: Humanity's Last Exam, dataset
cais/hle.
Subsets
Group
Languages
How the target-language text was produced
Source
en
Original English questions and answers.
Machine-translated and expert-verified / revised
zh, ja, ko, fr, th
Machine translation reviewed by a human… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HLE-BioMedX.details_lighteternal__Llama3-merge-biomed-8bBiomedCoOp
Biomedical Few-shot Image Classification for Vision-Language Models
Overview
Recent advancements in vision-language models (VLMs), such as CLIP, have demonstrated substantial success in self-supervised representation learning for vision tasks.
However, effectively adapting VLMs to downstream applications remains challenging, as their accuracy often depends on time-intensive and expertise-demanding prompt
engineering, while full model fine-tuning is… See the full description on the dataset page: https://huggingface.co/datasets/TahaKoleilat/BiomedCoOp.biomed-fr
Dataset Card for "biomed-fr"
More Information needed
gliner-biomed-pre-training
GLiNER-BioMed pre-training dataset
This dataset, used for the pre-training stage of the GLiNER-BioMed models, was introduced in the paper GLiNER-BioMed: a suite of efficient models for open biomedical named entity recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@article{yazdani2026gliner,
author = {Yazdani, Anthony and Stepanov, Ihor and Teodoro, Douglas},
title = {{GLiNER-BioMed}: a suite of efficient… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-pre-training.biomedical_lectures_eng_v2
Vidore Benchmark 2 - MIT Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to MIT biology courses. It includes a curated set of documents, queries, relevance judgments (qrels), and page images.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_eng_v2.Pile-NER-biomed-IOBRoP_biomedical
RoP — Biomedical Reference of Parameters
One CDE set to rule them all...
📋 Interested in help using RoP? → Submit Interest Form ← Start here
Executive Summary
RoP decreases activation energy in biomedical research by accelerating progress past the biggest bottlenecks ... data wrangling, interoperability and AI-readiness.
Right now, multi-cohort studies spend months (often years) manually mapping variables—PPMI calls it "MoCA Total Score," NACC calls it… See the full description on the dataset page: https://huggingface.co/datasets/DataTecnica/RoP_biomedical.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.multi_biomed_french_edu3_20250710gliner-biomed-post-training
GLiNER-BioMed post-training dataset
This dataset, used for the post-training stage of the GLiNER-BioMed models, was introduced in the paper GLiNER-BioMed: a suite of efficient models for open biomedical named entity recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@article{yazdani2026gliner,
author = {Yazdani, Anthony and Stepanov, Ihor and Teodoro, Douglas},
title = {{GLiNER-BioMed}: a suite of efficient… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-post-training.gliner-biomed-curated-corpus
GLiNER-BioMed curated corpus
Unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@misc{yazdani2025glinerbiomedsuiteefficientmodels,
title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition},
author={Anthony Yazdani and Ihor Stepanov and Douglas Teodoro}… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-curated-corpus.multi_biomed_french_prefixed_20250710gliner-biomed-balanced-curated-corpus
GLiNER-BioMed balanced curated corpus
Balanced, unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@misc{yazdani2025glinerbiomedsuiteefficientmodels,
title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition},
author={Anthony Yazdani and Ihor Stepanov and Douglas… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-balanced-curated-corpus.BioMedFlickr
BioMedFlickr
BioMedFlickr is a biomedical image–caption retrieval benchmark built from
public Flickr pathology / microscopy albums. Each example is a single image
paired with a cleaned English caption. Images are stored as original JPEGs inside parquet (~865 MB download). This dataset is the retrieval benchmark used in BIOMEDICA (Lozano et al.), published at CVPR 2025: https://cvpr.thecvf.com/virtual/2025/poster/33761
Load
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/BioMedFlickr.BiomedSQLBiomedSQL
GitHub
Paper
Dataset Summary
BiomedSQL is a text-to-SQL benchmark designed to evaluate Large Language Models (LLMs) on scientific tabular reasoning tasks. It consists of curated question-SQL query-answer triples covering a variety of biomedical and SQL reasoning types. The benchmark challenges models to apply implicit
scientific criteria rather than simply translating syntax.
Repository Organization
benchmark_data: contains the question-SQL query-answer triples.
Please note that you… See the full description on the dataset page: https://huggingface.co/datasets/NIH-CARD/BiomedSQL.biomed_NER
Biomed NER
This dataset consists of 4,840 manually annotated text records drawn from PubMed abstracts, drug descriptions from the FDA, and patent abstracts. All entities are continuous, and there are no nested entities.
Dataset composition
The dataset contains 4,840 annotated text records distributed across three sources:
Source
Approx. records
Purpose
PubMed abstracts
~4,300
Core biomedical content
FDA drug descriptions
~430
Pharmaceutical text with dense… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/biomed_NER.
