datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genomics-long-range-benchmarkDataset for benchmark of genomic deep learning models.synthetic-protein-folding-matrices
Dataset Card for CGSC Synthetic Protein Folding Matrices
Dataset Summary
This is the secondary repository for the CGSC, strictly dedicated to archiving high-resolution, uncompressed molecular dynamics (MD) trajectories. The dataset contains continuous temporal matrices representing synthetic protein folding simulations at sub-angstrom resolution.
Simulating atomic interactions over time generates colossal amounts of data. To preserve the micro-fluctuations and raw… See the full description on the dataset page: https://huggingface.co/datasets/comp-genomics-consortium/synthetic-protein-folding-matrices.raw-transcriptome-unmapped-reads
Dataset Card for CGSC Raw Transcriptome Unmapped Reads
Dataset Summary
This repository acts as the primary cold-storage for unmapped, raw sequencing outputs generated during the Q2 2026 Synthetic Bio-Arrays trials. The dataset comprises massive, uncompressed binary blobs that represent pre-alignment genomic data directly from the sequencing hardware.
Because these files bypass standard alignment and compression algorithms (such as BAM/CRAM conversion) to preserve… See the full description on the dataset page: https://huggingface.co/datasets/comp-genomics-consortium/raw-transcriptome-unmapped-reads.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.msr_genomics_kbcompThe database is derived from the NCI PID Pathway Interaction Database, and the textual mentions are extracted from cooccurring pairs of genes in PubMed abstracts, processed and annotated by Literome (Poon et al. 2014). This dataset was used in the paper “Compositional Learning of Embeddings for Relation Paths in Knowledge Bases and Text” (Toutanova, Lin, Yih, Poon, and Quirk, 2016).snpedia
SNPedia SQLite Cache
Pre-built SQLite cache of SNPedia genotype annotations for use in local genomic analysis tools.
What's in the file
snpedia.sqlite.gz is a gzipped SQLite database containing SNPedia's community-curated genotype annotations, scraped in compliance with SNPedia's robots.txt.
Stats
~175 MB uncompressed
~104,720 genotype entries
Indexed on rsID for fast lookup
Source and license
Source: SNPedia (community-curated… See the full description on the dataset page: https://huggingface.co/datasets/genomics-commons/snpedia.greengenes
Greengenes Dataset (modified for deeptaxa)
This dataset contains 16S rRNA gene sequences with hierarchical taxonomic annotations, designed for training and evaluating models like DeepTaxa. It is a processed version of the Greengenes database, widely used in microbiome research.
Dataset Details
The dataset includes the following files:
File Name
Type
Number of Sequences
Size
gg_2024_09_training.fna.gz
FASTA (sequences)
277,336
~96.4 MB… See the full description on the dataset page: https://huggingface.co/datasets/systems-genomics-lab/greengenes.applied-genomicsautocsf-genomics
AutoCSF genomics datasets
This artifact contains the original public FASTA/FASTQ inputs and deterministic
processed tables used by the AutoCSF genomics experiments. Each processed file
is a Zstandard-compressed, headerless UTF-8 TSV containing
15-mer<TAB>occurrence-count, sorted bytewise by k-mer.
The bundled dense 2-bit counter counts overlapping forward-strand 15-mers,
normalizes ASCII acgt to uppercase, keeps counts of one, and omits windows
containing other bases.
manifests/… See the full description on the dataset page: https://huggingface.co/datasets/detorresramos/autocsf-genomics.Synthetic-cancer-clinical-genomics
Synthetic Cancer Clinical Genomics Dataset
This repository contains multi-modal synthetic clinical-genomic data designed for oncology tracking, therapeutic progression analysis, and financial cost-modeling. The dataset is organized relationally under four core entities: Patients, Genomic Biomarkers, Clinical Encounters, and Financial Claims.
Dataset Structure
The dataset is provided as a unified JSON file containing four arrays of structured objects.
{
"patients": [...]… See the full description on the dataset page: https://huggingface.co/datasets/AnodeAI/Synthetic-cancer-clinical-genomics.cadd-scores
CADD v1.7 Scores (SQLite Cache)
Pre-built SQLite cache of CADD v1.7 PHRED-scaled variant deleteriousness scores, filtered to positions present in commonly used genomic databases.
What's in the file
cadd.sqlite.gz is a gzipped SQLite database containing CADD PHRED scores for SNV and indel variants. Indexed on (chromosome, position, alt) for fast coordinate lookup.
Source and license
Source: CADD v1.7 (Rentzsch et al. 2021, Schubach et al. 2024)… See the full description on the dataset page: https://huggingface.co/datasets/genomics-commons/cadd-scores.africa-synth-cancer-breast-cancer-genomics-ssa-all
Breast Cancer Genomics Synthetic Dataset (Sub-Saharan Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-breast-cancer-genomics-ssa-all.BioBench-Genomics
BioBench: A Bioinformatics Code Reasoning Benchmark
📑 Paper | 🌐 Project Page | 💾 Released Resources | 📦 Repo
This is the synthesized BioBench-Genomics dataset, designed for training LLMs on bioinformatics reasoning tasks.
Dataset
Dataset
Link
BioBench-Genomics
🤗
Please also check the raw data after our processing if you are interested:… See the full description on the dataset page: https://huggingface.co/datasets/toolevalxm/BioBench-Genomics.ocr-playground
clean.py
Dataset Summary
A finance dataset with pointcloud text modality, stored in lmdb format.
Preprocessing & Augmentation
Preprocessing: minimal
Augmentation: mixup cutmix
Splits & Sampling
Split strategy: random 90 10
Sampling: hard negative
Quality & Labeling
Quality filtering: lenient
Labeling: pseudo label
Files
clean.py — main artifact of this repository
License
See the… See the full description on the dataset page: https://huggingface.co/datasets/Jiangnan-genomics1/ocr-playground.3d-genomics-predictiongenomics-featurescontext-population-generalization-genomics-v01
Dataset
ClarusC64/context-population-generalization-genomics-v01
This dataset tests one capability.
Can a model keep genetic claims inside the population and context they were measured in.
Core rule
Genomic findings are population bound.
A claim must respect
ancestry
cohort design
sample context
transfer limits
What is true in one populationdoes not automatically hold in another.
Canonical labels
WITHIN_SCOPE
OUT_OF_SCOPE
Files… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/context-population-generalization-genomics-v01.uncertainty-incompleteness-functional-unknowns-genomics-v01
Dataset
ClarusC64/uncertainty-incompleteness-functional-unknowns-genomics-v01
This dataset tests one capability.
Can a model resist inventing biological function when evidence is incomplete.
Core rule
Genomics contains large unknowns.
A claim must respect
incomplete annotation
context specific regulation
limits of prediction
absence of functional validation
Prediction is not proof.
Annotation is not mechanism.
Expression is not causation.
Canonical labels… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/uncertainty-incompleteness-functional-unknowns-genomics-v01.africa-synth-cancer-cancer-genomics-molecular-africa-all
Cancer Genomics Molecular Africa | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-cancer-genomics-molecular-africa-all.comparative-genomics.GCF_005280335.1_ASM528033v1_genomiccausal-inference-variant-interpretation-genomics-v01
Dataset
ClarusC64/causal-inference-variant-interpretation-genomics-v01
This dataset tests one capability.
Can a model distinguish association from causation when interpreting genetic variants.
Core rule
Genomic evidence has tiers.
A claim must respect
evidence strength
effect size
penetrance
inheritance logic
Association does not equal causation.
Risk does not equal destiny.
Uncertain does not equal pathogenic.
Canonical labels
WITHIN_SCOPE… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/causal-inference-variant-interpretation-genomics-v01.highwire_trec-genomics-2007
Dataset Card for highwire/trec-genomics-2007
The highwire/trec-genomics-2007 dataset, provided by the ir-datasets package.
For more information about the dataset, see the documentation.
Data
This dataset provides:
queries (i.e., topics); count=36
qrels: (relevance assessments); count=35,996
Usage
from datasets import load_dataset
queries = load_dataset('irds/highwire_trec-genomics-2007', 'queries')
for record in queries:
record # {'query_id': ...… See the full description on the dataset page: https://huggingface.co/datasets/irds/highwire_trec-genomics-2007.comparative-genomics.GCF_023573625.1_ASM2357362v1_genomicgenomics
ClinVar Conflicting: https://www.kaggle.com/datasets/kevinarvai/clinvar-conflicting
VEP-Benchmark: https://github.com/DSIMB/VEP-Benchmark/tree/main
BMC_Medical_Genomics_Peer_Reviews
BMC Medical Genomics Peer Reviews
This dataset contains peer reviews from BMC, standardized to match the format of the pawin205/PeerRT dataset.
Dataset Structure
Each record contains the following attributes:
relative_rank: Default value (0).
win_prob: Default value (0.0).
title: Title of the paper.
abstract: Abstract of the paper.
full_text: Full text of the paper (or review text if unavailable).
review: The peer review text.
source: Source of the data ('BMC').… See the full description on the dataset page: https://huggingface.co/datasets/JerMa88/BMC_Medical_Genomics_Peer_Reviews.comparative-genomics.GCF_020097155.1_ASM2009715v1_genomicMedical_genomics
ANODE: Synthetic NSCLC EHR & Genomics Dataset
Title: ANODE-OncoGen High-Fidelity Synthetic Records
This dataset consists of a high-density .jsonl collection of synthetic patient records for Non-Small Cell Lung Cancer (NSCLC), integrating clinical history with deep genomic markers.
📋 Full Feature Specifications
The dataset includes the following features for every patient record:
1. Administrative & Metadata
patient_id: Unique identifier for the… See the full description on the dataset page: https://huggingface.co/datasets/AnodeAI/Medical_genomics.comparative-genomics.GCF_003691675.1_ASM369167v1_genomiccomparative-genomics.GCF_002008305.4_ASM200830v4_genomicgrls-genomics
GRLS Genotyping — Axiom Canine Array
Whole-genome SNP-array genotyping of golden retrievers from the Morris Animal Foundation Golden Retriever Lifetime Study (GRLS), mirrored from the public AWS Open Data registry (s3://mafgrlsgenome). Part of the golden-retrievers collection.
This is the openly-licensed genomics slice only. Phenotypes, clinical, behavioral, and biospecimen data live in the GRLS Data Commons and require an institutional data-use agreement — see "Linking… See the full description on the dataset page: https://huggingface.co/datasets/golden-retrievers/grls-genomics.
