datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gliner-biomed-curated-corpus
GLiNER-BioMed curated corpus
Unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@misc{yazdani2025glinerbiomedsuiteefficientmodels,
title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition},
author={Anthony Yazdani and Ihor Stepanov and Douglas Teodoro}… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-curated-corpus.gliner-biomed-balanced-curated-corpus
GLiNER-BioMed balanced curated corpus
Balanced, unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@misc{yazdani2025glinerbiomedsuiteefficientmodels,
title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition},
author={Anthony Yazdani and Ihor Stepanov and Douglas… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-balanced-curated-corpus.BiomedSQLBiomedSQL
GitHub
Paper
Dataset Summary
BiomedSQL is a text-to-SQL benchmark designed to evaluate Large Language Models (LLMs) on scientific tabular reasoning tasks. It consists of curated question-SQL query-answer triples covering a variety of biomedical and SQL reasoning types. The benchmark challenges models to apply implicit
scientific criteria rather than simply translating syntax.
Repository Organization
benchmark_data: contains the question-SQL query-answer triples.
Please note that you… See the full description on the dataset page: https://huggingface.co/datasets/NIH-CARD/BiomedSQL.BioMedGraphica
BioMedGraphica
BioMedGraphica is an all-in-one platform for biomedical data integration and knowledge graph generation. It harmonizes fragmented biomedical datasets into a unified, graph AI-ready resource that facilitates precision medicine, therapeutic target discovery, and integrative biomedical AI research.
Developed using data from 43 biomedical databases, BioMedGraphica integrates:
11 entity types
30 relation types… See the full description on the dataset page: https://huggingface.co/datasets/FuhaiLiAiLab/BioMedGraphica.bioleaflets-biomedical-ner
Dataset Card for BioLeaflets Dataset
Dataset Summary
BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website.
Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately.
This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.biomedical-topic-categorizationpredator-biomedical
PREDATOR Biomedical Dataset
50K+ curated biomedical abstracts from PubMed/EuropePMC
Curated biomedical research data extracted from PubMed, EuropePMC, and clinical trial databases. Each entry includes DOI/PMID, title, source, domain classification, commercial value score, and quality assessment.
Fields
Column
Description
id
DOI or PMID identifier
title
Article title or abstract summary
source
Data source (EuropePMC, PubMed, ClinicalTrials.gov)… See the full description on the dataset page: https://huggingface.co/datasets/iservice/predator-biomedical.BiomedKeyRefine
BiomedKeyRefine: Refined Biomedical Keyword Extraction Benchmark
Version 1.0Last updated: February 2026Author: Christian Hoang (@christhoang04)License: MIT
🌟 Dataset Summary
BiomedKeyRefine is a high-quality, refined subset of the PubMedAKE benchmark (CIKM 2022), specifically curated for biomedical keyword extraction research.
The original PubMedAKE contains 843k+ articles with author-assigned keywords (extractive + abstractive). This version:
Filters for… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/BiomedKeyRefine.biomedical_cpgQA
Dataset Card for the Biomedical Domain
Dataset Summary
This dataset was obtain through github (https://github.com/mmahbub/cpgQA/blob/main/dataset/cpgQA-v1.0.csv?plain=1) to Huggin Face for easier access while fine tuning.
Languages
English (en)
Dataset Structure
The dataset is in a CSV format, with each row representing a single review. The following columns are included:
Title: Categorises the QA.
Context: Gives a context of the QA.
Question: The… See the full description on the dataset page: https://huggingface.co/datasets/chloecchng/biomedical_cpgQA.biomedica_csvsbiomedical-topic-categorization-validationBiomedicalQuestionAnsweringDatasetintent-recognition-biomedicalsource
biomedical-datasetbiomedica-filelistBioMedJImpact
BioMedJImpact
Dataset Summary
This release packages two journal-level tables from the BioMedJImpact project for reviewer-friendly hosting on the Hugging Face Hub:
BioMedJImpact-2019.csv: 1,367 journals with metrics spanning 2016-2019.
BioMedJImpact-2023.csv: 2,685 journals with metrics spanning 2020-2023.
Each row represents one journal. Columns include journal identifiers, publisher metadata, impact factors, quartiles, citation/reference counts, publication… See the full description on the dataset page: https://huggingface.co/datasets/RuiyuWangRWAN388/BioMedJImpact.wmt22-biomedBiomedical_Entity_Relation_ExtractionBiomedTLDRA large sample of researcher-authored summaries from scientific papers, which leverages the common practice of including authors' comments alongside bibliography items.
Paper: https://arxiv.org/abs/2512.23206
biomedical_novelty_claim_corpus
Biomedical Novelty Claim Corpus (Sample)
This repository contains a sample of the Biomedical Novelty Claim Corpus, developed as part of our study on categorizing and analyzing novelty statements in biomedical research publications.
🔍 Overview
The dataset includes structured annotations of author-claimed novelty statements extracted from biomedical literature. It supports research in:
Scientific novelty classification
Biomedical text mining
Evidence-based discovery… See the full description on the dataset page: https://huggingface.co/datasets/clinicalnlplab/biomedical_novelty_claim_corpus.wmt16_biomed_testbiomedical_drug_protein_interaction_benchmarkswmt16_biomed_gold
