datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bioinformatics
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Bioinformatics.Bioinformatics
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/Bioinformatics.mst_gesynthetic-hla-frequencies
Synthetic HLA haplotype frequencies
This data is synthetic. It does not describe any real population, cohort, registry or
individual. Do not use it for clinical, epidemiological or population-genetics conclusions.
A generated five-locus (A~C~B~DRB1~DQB1) HLA haplotype frequency dataset, built as a
license-free test fixture for tools that consume haplotype frequency data — format
validators, converters, QC and diversity metrics, and publishing pipelines. Real frequency
datasets… See the full description on the dataset page: https://huggingface.co/datasets/nmdp-bioinformatics/synthetic-hla-frequencies.stackexchange_bioinformaticsbioinformatics-qa-dataset
Bioinformatics QA Dataset
A curated question-answer dataset for bioinformatics and computational biology model training.
Summary
Total examples: 5880
Unique topics: 65
Columns: id, topic, question, answer
Source format: CSV converted to Hugging Face Dataset
Intended Use
This dataset is intended for:
Instruction tuning and domain adaptation for biomedical and bioinformatics LLMs
QA benchmarking in life-science terminology
Prompt-response training pipelines… See the full description on the dataset page: https://huggingface.co/datasets/yashm/bioinformatics-qa-dataset.Bioinformatics-Toolkit-HYA
Bioinformatics Toolkit v1.0
Bioinformatics is plagued by software dependency drama. The Bioinformatics Toolkit makes it easy to set up a computational software environment. This allows researchers and students to focus on science rather than software troubleshooting.
It's a pre-built bioinformatics env for macOS that bundles the Python 3.12 interpreter and the uv binary. Packages are auto downloaded during the first run. The app is launched with a double-click.
This is a Hybrid App… See the full description on the dataset page: https://huggingface.co/datasets/vbookshelf/Bioinformatics-Toolkit-HYA.pubmed-bioinformatics-abstracts
PubMed Bioinformatics Abstracts (2020–2026)
A curated collection of 174,154 bioinformatics journal article abstracts from PubMed,
covering publications from 2020 to 2026.
Curated by: Dr. Yash M Gupta
Intended Use
Fine-tuning large language models (LLMs) on biomedical / bioinformatics domain text.
Dataset Structure
Column
Type
Description
pmid
string
PubMed ID
title
string
Article title
abstract
string
Full abstract text
year
string
Publication… See the full description on the dataset page: https://huggingface.co/datasets/yashm/pubmed-bioinformatics-abstracts.pubmed-bioinformatics-abstracts_all_years
PubMed Bioinformatics Abstracts — All Years
A curated dataset of bioinformatics research abstracts fetched from PubMed via the NCBI Entrez API, covering publications from 1990 through 2026. Designed for LLM fine-tuning, biomedical NLP, and scientific text analysis.
Curated by: Dr. Yash M Gupta
Pipeline: yashmgupta/pubmed-bioinformatics-abstracts
Intended Use
Fine-tuning large language models (LLMs) on biomedical / bioinformatics domain text.
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/yashm/pubmed-bioinformatics-abstracts_all_years.
