datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpn-star-p-uniform-v1-cds
marin-dna/gpn-star-p-uniform-v1-cds
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-cds.zoonomia-v1-v4_ccre_noexon
bolinas-dna/zoonomia-v1-v4_ccre_noexon
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is the v4 ccre_non_promoter arm with every window that overlaps any other functional element (CDS / 3′UTR / ncRNA exon / TSS+5′UTR) removed — i.e. windows whose functional content… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon.gpn-star-p-uniform-v1-enhancer-arm-a
marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.gpn-star-p-uniform-v1-background
marin-dna/gpn-star-p-uniform-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-background.phylop-uniform-v1-cds
marin-dna/phylop-uniform-v1-cds
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-cds.zoonomia-v1-v4_ccre_noexon_enhancer
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.functional-cds
marin-dna/functional-cds
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains.
This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
For the 28 non-mammalian targets, the stable ucsc_multiz100way… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-cds.zoonomia-v1-v3_cds
bolinas-dna/zoonomia-v1-v3_cds
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled cds by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (cds)
Coding sequence — Ensembl r115 CDS features (get_cds). Highest-priority class: any anchor with overlap on a CDS feature (and union-of-functional fraction ≥ 0.20 across all five labels) is labelled cds, regardless… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_cds.DNA_Gen
Citation
Please cite our work using the bibtex below:
BibTeX:
@article{su2025language,
title={Language Models for Controllable DNA Sequence Design},
author={Su, Xingyu and Li, Xiner and Lin, Yuchao and Xie, Ziqian and Zhi, Degui and Ji, Shuiwang},
journal={arXiv preprint arXiv:2507.19523},
year={2025}
}
vertebrate-v1-issue473-fullwindow-cds-random-val
marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val
CDS full-window vertebrate projection sequences for the issue #473 random
validation control. The source is the immutable issue #417 accepted-sequence
table.
The split uniformly samples 16,384 original-orientation CDS rows
without replacement using seed 42. Sampling occurs before
reverse-complement augmentation. Selected rows are removed from training;
reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.human_ref_dna
Dataset Card for "human_ref_dna"
More Information needed
zoonomia-v1-v3_ccre_non_promoter
bolinas-dna/zoonomia-v1-v3_ccre_non_promoter
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ccre_non_promoter by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (ccre_non_promoter)
ENCODE cCRE V4 non-promoter classes — cre_class != "PLS" (so: dELS, pELS, CA, CA-CTCF, CA-TF, CA-H3K4me3, TF), extended by 500 bp on each side. PLS is excluded because… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ccre_non_promoter.functional-enhancer
marin-dna/functional-enhancer
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
For the 28 non-mammalian targets, the stable… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-enhancer.zoonomia-v1-v1
zoonomia-v1-v1
255 bp human-anchored windows, conservation-filtered (phyloP_447m, proportion_conserved >= 0.20), projected onto 108 family-deduped Zoonomia 447-mammalian assemblies via halLiftover, midpoint-resized to 255 bp, reverse-complement-augmented. Single train split.
Produced by the zoonomia_projection_dataset pipeline in Open-Athena/bolinas-dna — permalinked at the exact code that built this dataset: snakemake/zoonomia_projection_dataset @ 7ff07cd (PR #158).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v1.zoonomia-v1-v3_ncrna_exon
bolinas-dna/zoonomia-v1-v3_ncrna_exon
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ncrna_exon by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (ncrna_exon)
Non-coding-RNA exon — every Ensembl r115 exon that is not part of a protein-coding transcript (get_exons(ann) − get_ensembl_protein_coding_exons(ann)). No biotype or quality filter, so this… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ncrna_exon.gpn-star-p-uniform-v1-utr3
marin-dna/gpn-star-p-uniform-v1-utr3
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the utr3 region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-utr3.vertebrate-v1-all
marin-dna/vertebrate-v1-all
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the all region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source non-repeat-masked… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-all.phylop-uniform-v1-enhancer-arm-a
marin-dna/phylop-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.vertebrate-v1-issue473-center1-cds
marin-dna/vertebrate-v1-issue473-center1-cds
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
cds cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed #417 protein-coding CDS anchor catalog.
The source projection was produced by the… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-cds.phylop-uniform-v1-ncrna-exon
marin-dna/phylop-uniform-v1-ncrna-exon
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the ncrna_exon region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-ncrna-exon.gpn-star-p-uniform-v1-ncrna-exon
marin-dna/gpn-star-p-uniform-v1-ncrna-exon
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the ncrna_exon region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-ncrna-exon.v100-ds-mmlugpn-star-p-uniform-v1-tss-utr5
marin-dna/gpn-star-p-uniform-v1-tss-utr5
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the tss_region_and_utr5 region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-tss-utr5.zoonomia-v1-v4_ccre_non_promoter-order
bolinas-dna/zoonomia-v1-v4_ccre_non_promoter-order
The bolinas-dna/zoonomia-v1-v4_ccre_non_promoter cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_non_promoter-order.prs-percentilesalphagenome_avi
AlphaGenome AVI scores, re-encoded
AlphaGenome's Variant Impact (AVI) scores for 8,812,917,339 SNVs on GRCh38, re-encoded from
the 88.5 GB published tabix TSV into ~34 GB of parquet by
just-dna-enricher.
This is a re-encoding, not a re-analysis. No score is changed, recomputed or filtered.
What is in it
data/alphagenome_avi-<contig>.parquet
chrom, pos (1-based VCF), ref, alt, raw_score_e5
avi_knots.parquet
the PHRED reconstruction curve — not… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/alphagenome_avi.phylop-uniform-v1-tss-utr5
marin-dna/phylop-uniform-v1-tss-utr5
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the tss_region_and_utr5 region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-tss-utr5.zoonomia-v1-v3_utr3
bolinas-dna/zoonomia-v1-v3_utr3
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled utr3 by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (utr3)
3' untranslated region — Ensembl r115 protein-coding transcripts' 3' UTR (get_ensembl_3_prime_utr, filtered to transcript_biotype "protein_coding"). Second-priority class: wins over ncrna_exon… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_utr3.zoonomia-v1-v4_cds-order
bolinas-dna/zoonomia-v1-v4_cds-order
The bolinas-dna/zoonomia-v1-v4_cds cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors and same per-window… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_cds-order.vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the full_window policy.
The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.
