CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01InstaDeepAI /genomics-long-range-benchmarkDataset for benchmark of genomic deep learning models.15 likes1.8k downloads2y agoHugging Face02comp-genomics-consortium /synthetic-protein-folding-matrices Dataset Card for CGSC Synthetic Protein Folding Matrices Dataset Summary This is the secondary repository for the CGSC, strictly dedicated to archiving high-resolution, uncompressed molecular dynamics (MD) trajectories. The dataset contains continuous temporal matrices representing synthetic protein folding simulations at sub-angstrom resolution. Simulating atomic interactions over time generates colossal amounts of data. To preserve the micro-fluctuations and raw… See the full description on the dataset page: https://huggingface.co/datasets/comp-genomics-consortium/synthetic-protein-folding-matrices.time-series-forecasting100K<n<1M0 likes1.4k downloads2mo agoHugging Face03comp-genomics-consortium /raw-transcriptome-unmapped-reads Dataset Card for CGSC Raw Transcriptome Unmapped Reads Dataset Summary This repository acts as the primary cold-storage for unmapped, raw sequencing outputs generated during the Q2 2026 Synthetic Bio-Arrays trials. The dataset comprises massive, uncompressed binary blobs that represent pre-alignment genomic data directly from the sequencing hardware. Because these files bypass standard alignment and compression algorithms (such as BAM/CRAM conversion) to preserve… See the full description on the dataset page: https://huggingface.co/datasets/comp-genomics-consortium/raw-transcriptome-unmapped-reads.tabular-classification100K<n<1M1 likes1.2k downloads3mo agoHugging Face04simpleG2023 /chinese-biomedicine-and-genomics-open-intelligence 🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes240 downloads1d agoHugging Face05microsoft /msr_genomics_kbcompThe database is derived from the NCI PID Pathway Interaction Database, and the textual mentions are extracted from cooccurring pairs of genes in PubMed abstracts, processed and annotated by Literome (Poon et al. 2014). This dataset was used in the paper “Compositional Learning of Embeddings for Relation Paths in Knowledge Bases and Text” (Toutanova, Lin, Yih, Poon, and Quirk, 2016).other10K<n<100K3 likes167 downloads3y agoHugging Face06genomics-commons /snpedia SNPedia SQLite Cache Pre-built SQLite cache of SNPedia genotype annotations for use in local genomic analysis tools. What's in the file snpedia.sqlite.gz is a gzipped SQLite database containing SNPedia's community-curated genotype annotations, scraped in compliance with SNPedia's robots.txt. Stats ~175 MB uncompressed ~104,720 genotype entries Indexed on rsID for fast lookup Source and license Source: SNPedia (community-curated… See the full description on the dataset page: https://huggingface.co/datasets/genomics-commons/snpedia.100K<n<1M0 likes99 downloads4mo agoHugging Face07systems-genomics-lab /greengenes Greengenes Dataset (modified for deeptaxa) This dataset contains 16S rRNA gene sequences with hierarchical taxonomic annotations, designed for training and evaluating models like DeepTaxa. It is a processed version of the Greengenes database, widely used in microbiome research. Dataset Details The dataset includes the following files: File Name Type Number of Sequences Size gg_2024_09_training.fna.gz FASTA (sequences) 277,336 ~96.4 MB… See the full description on the dataset page: https://huggingface.co/datasets/systems-genomics-lab/greengenes.texttext-classification100K<n<1M0 likes97 downloads1y agoHugging Face08funlab /applied-genomicstabularn<1K0 likes52 downloads2y agoHugging Face09detorresramos /autocsf-genomics AutoCSF genomics datasets This artifact contains the original public FASTA/FASTQ inputs and deterministic processed tables used by the AutoCSF genomics experiments. Each processed file is a Zstandard-compressed, headerless UTF-8 TSV containing 15-mer<TAB>occurrence-count, sorted bytewise by k-mer. The bundled dense 2-bit counter counts overlapping forward-strand 15-mers, normalizes ASCII acgt to uppercase, keeps counts of one, and omits windows containing other bases. manifests/… See the full description on the dataset page: https://huggingface.co/datasets/detorresramos/autocsf-genomics.tabular-classification0 likes48 downloads1mo agoHugging Face10AnodeAI /Synthetic-cancer-clinical-genomics Synthetic Cancer Clinical Genomics Dataset This repository contains multi-modal synthetic clinical-genomic data designed for oncology tracking, therapeutic progression analysis, and financial cost-modeling. The dataset is organized relationally under four core entities: Patients, Genomic Biomarkers, Clinical Encounters, and Financial Claims. Dataset Structure The dataset is provided as a unified JSON file containing four arrays of structured objects. { "patients": [...]… See the full description on the dataset page: https://huggingface.co/datasets/AnodeAI/Synthetic-cancer-clinical-genomics.texttabular-classification10K<n<100K2 likes42 downloads4mo agoHugging Face11genomics-commons /cadd-scores CADD v1.7 Scores (SQLite Cache) Pre-built SQLite cache of CADD v1.7 PHRED-scaled variant deleteriousness scores, filtered to positions present in commonly used genomic databases. What's in the file cadd.sqlite.gz is a gzipped SQLite database containing CADD PHRED scores for SNV and indel variants. Indexed on (chromosome, position, alt) for fast coordinate lookup. Source and license Source: CADD v1.7 (Rentzsch et al. 2021, Schubach et al. 2024)… See the full description on the dataset page: https://huggingface.co/datasets/genomics-commons/cadd-scores.100M<n<1B0 likes39 downloads3mo agoHugging Face12electricsheepafrica /africa-synth-cancer-breast-cancer-genomics-ssa-all Breast Cancer Genomics Synthetic Dataset (Sub-Saharan Africa) | Africa (Electric Sheep Africa metadata inventory) Size category: 100K<n<1M - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-breast-cancer-genomics-ssa-all.tabulartabular-classification100K<n<1M1 likes38 downloads1mo agoHugging Face13toolevalxm /BioBench-Genomics BioBench: A Bioinformatics Code Reasoning Benchmark 📑 Paper    |    🌐 Project Page    |    💾 Released Resources    |    📦 Repo This is the synthesized BioBench-Genomics dataset, designed for training LLMs on bioinformatics reasoning tasks. Dataset Dataset Link BioBench-Genomics 🤗 Please also check the raw data after our processing if you are interested:… See the full description on the dataset page: https://huggingface.co/datasets/toolevalxm/BioBench-Genomics.0 likes34 downloads7mo agoHugging Face14Jiangnan-genomics1 /ocr-playground clean.py Dataset Summary A finance dataset with pointcloud text modality, stored in lmdb format. Preprocessing & Augmentation Preprocessing: minimal Augmentation: mixup cutmix Splits & Sampling Split strategy: random 90 10 Sampling: hard negative Quality & Labeling Quality filtering: lenient Labeling: pseudo label Files clean.py — main artifact of this repository License See the… See the full description on the dataset page: https://huggingface.co/datasets/Jiangnan-genomics1/ocr-playground.0 likes33 downloads27d agoHugging Face15FatimaInstitute /3d-genomics-predictiongated0 likes29 downloads25d agoHugging Face16huseyincavus /genomics-features0 likes29 downloads4d agoHugging Face17ClarusC64 /context-population-generalization-genomics-v01 Dataset ClarusC64/context-population-generalization-genomics-v01 This dataset tests one capability. Can a model keep genetic claims inside the population and context they were measured in. Core rule Genomic findings are population bound. A claim must respect ancestry cohort design sample context transfer limits What is true in one populationdoes not automatically hold in another. Canonical labels WITHIN_SCOPE OUT_OF_SCOPE Files… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/context-population-generalization-genomics-v01.texttext-classificationn<1K0 likes28 downloads8mo agoHugging Face18ClarusC64 /uncertainty-incompleteness-functional-unknowns-genomics-v01 Dataset ClarusC64/uncertainty-incompleteness-functional-unknowns-genomics-v01 This dataset tests one capability. Can a model resist inventing biological function when evidence is incomplete. Core rule Genomics contains large unknowns. A claim must respect incomplete annotation context specific regulation limits of prediction absence of functional validation Prediction is not proof. Annotation is not mechanism. Expression is not causation. Canonical labels… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/uncertainty-incompleteness-functional-unknowns-genomics-v01.texttext-classificationn<1K0 likes26 downloads8mo agoHugging Face19electricsheepafrica /africa-synth-cancer-cancer-genomics-molecular-africa-all Cancer Genomics Molecular Africa | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-cancer-genomics-molecular-africa-all.tabulartabular-classification1K<n<10K0 likes24 downloads1mo agoHugging Face20entropic-digital /comparative-genomics.GCF_005280335.1_ASM528033v1_genomictextn<1K0 likes22 downloads10mo agoHugging Face21ClarusC64 /causal-inference-variant-interpretation-genomics-v01 Dataset ClarusC64/causal-inference-variant-interpretation-genomics-v01 This dataset tests one capability. Can a model distinguish association from causation when interpreting genetic variants. Core rule Genomic evidence has tiers. A claim must respect evidence strength effect size penetrance inheritance logic Association does not equal causation. Risk does not equal destiny. Uncertain does not equal pathogenic. Canonical labels WITHIN_SCOPE… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/causal-inference-variant-interpretation-genomics-v01.texttext-classificationn<1K0 likes22 downloads8mo agoHugging Face22irds /highwire_trec-genomics-2007 Dataset Card for highwire/trec-genomics-2007 The highwire/trec-genomics-2007 dataset, provided by the ir-datasets package. For more information about the dataset, see the documentation. Data This dataset provides: queries (i.e., topics); count=36 qrels: (relevance assessments); count=35,996 Usage from datasets import load_dataset queries = load_dataset('irds/highwire_trec-genomics-2007', 'queries') for record in queries: record # {'query_id': ...… See the full description on the dataset page: https://huggingface.co/datasets/irds/highwire_trec-genomics-2007.text-retrieval1 likes21 downloads4y agoHugging Face23entropic-digital /comparative-genomics.GCF_023573625.1_ASM2357362v1_genomictextn<1K0 likes21 downloads10mo agoHugging Face24huseyincavus /genomics ClinVar Conflicting: https://www.kaggle.com/datasets/kevinarvai/clinvar-conflicting VEP-Benchmark: https://github.com/DSIMB/VEP-Benchmark/tree/main tabular-classification0 likes20 downloads6mo agoHugging Face25JerMa88 /BMC_Medical_Genomics_Peer_Reviews BMC Medical Genomics Peer Reviews This dataset contains peer reviews from BMC, standardized to match the format of the pawin205/PeerRT dataset. Dataset Structure Each record contains the following attributes: relative_rank: Default value (0). win_prob: Default value (0.0). title: Title of the paper. abstract: Abstract of the paper. full_text: Full text of the paper (or review text if unavailable). review: The peer review text. source: Source of the data ('BMC').… See the full description on the dataset page: https://huggingface.co/datasets/JerMa88/BMC_Medical_Genomics_Peer_Reviews.tabular1K<n<10K0 likes19 downloads10mo agoHugging Face26entropic-digital /comparative-genomics.GCF_020097155.1_ASM2009715v1_genomictextn<1K0 likes19 downloads10mo agoHugging Face27AnodeAI /Medical_genomics ANODE: Synthetic NSCLC EHR & Genomics Dataset Title: ANODE-OncoGen High-Fidelity Synthetic Records This dataset consists of a high-density .jsonl collection of synthetic patient records for Non-Small Cell Lung Cancer (NSCLC), integrating clinical history with deep genomic markers. 📋 Full Feature Specifications The dataset includes the following features for every patient record: 1. Administrative & Metadata patient_id: Unique identifier for the… See the full description on the dataset page: https://huggingface.co/datasets/AnodeAI/Medical_genomics.tabulartabular-classification10K<n<100K1 likes17 downloads6mo agoHugging Face28entropic-digital /comparative-genomics.GCF_003691675.1_ASM369167v1_genomictextn<1K0 likes16 downloads10mo agoHugging Face29entropic-digital /comparative-genomics.GCF_002008305.4_ASM200830v4_genomictextn<1K0 likes15 downloads10mo agoHugging Face30golden-retrievers /grls-genomics GRLS Genotyping — Axiom Canine Array Whole-genome SNP-array genotyping of golden retrievers from the Morris Animal Foundation Golden Retriever Lifetime Study (GRLS), mirrored from the public AWS Open Data registry (s3://mafgrlsgenome). Part of the golden-retrievers collection. This is the openly-licensed genomics slice only. Phenotypes, clinical, behavioral, and biospecimen data live in the GRLS Data Commons and require an institutional data-use agreement — see "Linking… See the full description on the dataset page: https://huggingface.co/datasets/golden-retrievers/grls-genomics.1K<n<10K0 likes14 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.