datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpn-star-scores
GPN-Star genome-wide scores
Genome-wide, mutation-rate-calibrated GPN-Star constraint and variant scores for
eight score sets covering human, mouse, chicken, D. melanogaster,
C. elegans, and A. thaliana. Canonical scores are available as
chromosome-sharded Parquet. Hugging Face hosts 72 BigWigs; a multi-assembly
UCSC track hub references the 64 logo/LLR views.
Overview and quick links
Resource
Link
Files
Browse all dataset files
Default Dataset… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-scores.gpn-star-p-uniform-v1-cds
marin-dna/gpn-star-p-uniform-v1-cds
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-cds.gpn-animal-promoter-datasetgpn-star-p-uniform-v1-enhancer-arm-a
marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.gpn-msa-hg38-scores
GPN-MSA predictions for all possible SNPs in the human genome (~9 billion)
For more information check out our paper and repository.
Querying specific variants or genes
Install the latest tabix:In your current conda environment (might be slow):conda install -c bioconda -c conda-forge htslib=1.18
or in a new conda environment:conda create -n tabix -c bioconda -c conda-forge htslib=1.18
conda activate tabix
Query a specific region (e.g. BRCA1), from the remote file:… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-msa-hg38-scores.gpn_grassgpn-star-p-uniform-v1-background
marin-dna/gpn-star-p-uniform-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-background.gpn-star-p-uniform-v1-utr3
marin-dna/gpn-star-p-uniform-v1-utr3
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the utr3 region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-utr3.gpn-star-p-uniform-v1-tss-utr5
marin-dna/gpn-star-p-uniform-v1-tss-utr5
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the tss_region_and_utr5 region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-tss-utr5.gpn-star-p-uniform-v1-ncrna-exon
marin-dna/gpn-star-p-uniform-v1-ncrna-exon
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the ncrna_exon region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-ncrna-exon.gpn-msa-sapiens-dataset
Training windows for GPN-MSA-Sapiens
For more information check out our paper and repository.
Path in Snakemake:
results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001
gpn-msa-microglia-fullgpn-animal-promoter-checkpoints
GPN-Animal-Promoter intermediate checkpoints
# created using this command
hf upload-large-folder --repo-type=dataset songlab/gpn-animal-promoter-checkpoints results/checkpoints/mlm/v4_v6/512_64_1024
ucsc-tracks-gpn-arabidopsisgpn-animal-promoter-checkpoints-second-part
checkpoints
This model is a fine-tuned version of songlab/gpn-animal-promoter on the dataset dataset.
It achieves the following results on the evaluation set:
Loss: 1.1658
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:… See the full description on the dataset page: https://huggingface.co/datasets/gonzalobenegas/gpn-animal-promoter-checkpoints-second-part.gpn-msa-hg38-gene-essentiality-scores
GPN-MSA gene essentiality scores for the human genome
For more information check out our paper and repository.
gpn-animal-promoter-dataset-metadata
GPN promoter dataset metadata mapping
Commands
# Extract unique RefSeq IDs from GPN promoter dataset
python gpn_promoter_accession_extraction.py --output_dir=$(pwd) | tee gpn_promoter_accession_extraction.log 2>&1
# Fetch taxonomy information for RefSeq IDs
bash gpn_promoter_accession_mapping.sh unique_refseq_ids_train.txt refseq_taxonomy_results_train.tsv
bash gpn_promoter_accession_mapping.sh unique_refseq_ids_validation.txt refseq_taxonomy_results_validation.tsv
bash… See the full description on the dataset page: https://huggingface.co/datasets/eczech/gpn-animal-promoter-dataset-metadata.gpn-star-umap-regions
GPN-Star UMAP Regions
A set of 111,329 labeled 100 bp windows of the human genome (hg38/GRCh38), spanning
seven functional region classes (100 bp ≈ the median human coding-exon length). This is
the exact region set used for the GPN-Star embedding UMAP visualization in
Ye et al. 2025.
It can be used as:
a benchmark for genomic region / functional-element classification, and
a labeled region set for interpretation of genomic models (e.g. UMAP / probing of
sequence embeddings).… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-umap-regions.gpn_combined_random_uniform_10Mb_256gpn-animal-promoter-early-checkpoints
checkpoints
This model is a fine-tuned version of on the dataset dataset.
It achieves the following results on the evaluation set:
Loss: 1.2163
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 0.001… See the full description on the dataset page: https://huggingface.co/datasets/gonzalobenegas/gpn-animal-promoter-early-checkpoints.gpn-animal-promoter-datasetgpn_original_balanced_v1gpn_grass_balanced_v1gpn-msa-with-microgliacustom-pangenome-gpngpn_combined_random_uniform_10Mb_512gpn_combined_balanced_v1gpn-checkpointsgpn_grass_120kgpnBambooAll
