genomics
Datasets
All datasets matching “genomics”genomics-long-range-benchmarkDataset for benchmark of genomic deep learning models.synthetic-protein-folding-matrices
Dataset Card for CGSC Synthetic Protein Folding Matrices
Dataset Summary
This is the secondary repository for the CGSC, strictly dedicated to archiving high-resolution, uncompressed molecular dynamics (MD) trajectories. The dataset contains continuous temporal matrices representing synthetic protein folding simulations at sub-angstrom resolution.
Simulating atomic interactions over time generates colossal amounts of data. To preserve the micro-fluctuations and raw… See the full description on the dataset page: https://huggingface.co/datasets/comp-genomics-consortium/synthetic-protein-folding-matrices.raw-transcriptome-unmapped-reads
Dataset Card for CGSC Raw Transcriptome Unmapped Reads
Dataset Summary
This repository acts as the primary cold-storage for unmapped, raw sequencing outputs generated during the Q2 2026 Synthetic Bio-Arrays trials. The dataset comprises massive, uncompressed binary blobs that represent pre-alignment genomic data directly from the sequencing hardware.
Because these files bypass standard alignment and compression algorithms (such as BAM/CRAM conversion) to preserve… See the full description on the dataset page: https://huggingface.co/datasets/comp-genomics-consortium/raw-transcriptome-unmapped-reads.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.msr_genomics_kbcompThe database is derived from the NCI PID Pathway Interaction Database, and the textual mentions are extracted from cooccurring pairs of genes in PubMed abstracts, processed and annotated by Literome (Poon et al. 2014). This dataset was used in the paper “Compositional Learning of Embeddings for Relation Paths in Knowledge Bases and Text” (Toutanova, Lin, Yih, Poon, and Quirk, 2016).snpedia
SNPedia SQLite Cache
Pre-built SQLite cache of SNPedia genotype annotations for use in local genomic analysis tools.
What's in the file
snpedia.sqlite.gz is a gzipped SQLite database containing SNPedia's community-curated genotype annotations, scraped in compliance with SNPedia's robots.txt.
Stats
~175 MB uncompressed
~104,720 genotype entries
Indexed on rsID for fast lookup
Source and license
Source: SNPedia (community-curated… See the full description on the dataset page: https://huggingface.co/datasets/genomics-commons/snpedia.
