datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zoonomia-v1-v4_ccre_noexon
bolinas-dna/zoonomia-v1-v4_ccre_noexon
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is the v4 ccre_non_promoter arm with every window that overlaps any other functional element (CDS / 3′UTR / ncRNA exon / TSS+5′UTR) removed — i.e. windows whose functional content… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon.zoonomia-v1-v4_ccre_noexon_enhancer
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.zoonomia-v1-v3_cds
bolinas-dna/zoonomia-v1-v3_cds
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled cds by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (cds)
Coding sequence — Ensembl r115 CDS features (get_cds). Highest-priority class: any anchor with overlap on a CDS feature (and union-of-functional fraction ≥ 0.20 across all five labels) is labelled cds, regardless… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_cds.zoonomia-v1-v3_ccre_non_promoter
bolinas-dna/zoonomia-v1-v3_ccre_non_promoter
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ccre_non_promoter by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (ccre_non_promoter)
ENCODE cCRE V4 non-promoter classes — cre_class != "PLS" (so: dELS, pELS, CA, CA-CTCF, CA-TF, CA-H3K4me3, TF), extended by 500 bp on each side. PLS is excluded because… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ccre_non_promoter.zoonomia-v1-v1
zoonomia-v1-v1
255 bp human-anchored windows, conservation-filtered (phyloP_447m, proportion_conserved >= 0.20), projected onto 108 family-deduped Zoonomia 447-mammalian assemblies via halLiftover, midpoint-resized to 255 bp, reverse-complement-augmented. Single train split.
Produced by the zoonomia_projection_dataset pipeline in Open-Athena/bolinas-dna — permalinked at the exact code that built this dataset: snakemake/zoonomia_projection_dataset @ 7ff07cd (PR #158).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v1.zoonomia-v1-v3_ncrna_exon
bolinas-dna/zoonomia-v1-v3_ncrna_exon
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ncrna_exon by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (ncrna_exon)
Non-coding-RNA exon — every Ensembl r115 exon that is not part of a protein-coding transcript (get_exons(ann) − get_ensembl_protein_coding_exons(ann)). No biotype or quality filter, so this… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ncrna_exon.zoonomia-v1-v4_ccre_non_promoter-order
bolinas-dna/zoonomia-v1-v4_ccre_non_promoter-order
The bolinas-dna/zoonomia-v1-v4_ccre_non_promoter cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_non_promoter-order.zoonomia-v1-v4_ccre_non_promoter
bolinas-dna/zoonomia-v1-v4_ccre_non_promoter
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ccre_non_promoter by the v4 region labeler
(snakemake/zoonomia_projection_dataset pipeline,
commit 4729d06d576f).
v4 re-derives the v3 partition with the labeling scheme resolved in
issue #221: base-pair
priority + window majority, a protein-coding-only TSS band, and
ccre_flank=0. See the pipeline README's "v4… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_non_promoter.zoonomia-v1-v4_cds
bolinas-dna/zoonomia-v1-v4_cds
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled cds by the v4 region labeler
(snakemake/zoonomia_projection_dataset pipeline,
commit 4729d06d576f).
v4 re-derives the v3 partition with the labeling scheme resolved in
issue #221: base-pair
priority + window majority, a protein-coding-only TSS band, and
ccre_flank=0. See the pipeline README's "v4 region-type annotation"… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_cds.zoonomia-v1-v3_utr3
bolinas-dna/zoonomia-v1-v3_utr3
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled utr3 by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (utr3)
3' untranslated region — Ensembl r115 protein-coding transcripts' 3' UTR (get_ensembl_3_prime_utr, filtered to transcript_biotype "protein_coding"). Second-priority class: wins over ncrna_exon… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_utr3.zoonomia-v1-v4_cds-order
bolinas-dna/zoonomia-v1-v4_cds-order
The bolinas-dna/zoonomia-v1-v4_cds cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors and same per-window… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_cds-order.zoonomia-v1-v4_ncrna_exon
bolinas-dna/zoonomia-v1-v4_ncrna_exon
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ncrna_exon by the v4 region labeler
(snakemake/zoonomia_projection_dataset pipeline,
commit 4729d06d576f).
v4 re-derives the v3 partition with the labeling scheme resolved in
issue #221: base-pair
priority + window majority, a protein-coding-only TSS band, and
ccre_flank=0. See the pipeline README's "v4 region-type… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ncrna_exon.zoonomia-v1-v4_ccre_noexon_enhancer-order
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order
The bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order.zoonomia-v1-v3_tss_region_and_utr5
bolinas-dna/zoonomia-v1-v3_tss_region_and_utr5
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled tss_region_and_utr5 by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (tss_region_and_utr5)
TSS region and 5' UTR — (TSS ± 256 bp on every Ensembl r115 transcript) ∪ 5' UTR of every protein-coding transcript. One class instead of separate promoter + utr5… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_tss_region_and_utr5.zoonomia-v1-v4_bg
bolinas-dna/zoonomia-v1-v4_bg
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled background by the v4 region labeler
(snakemake/zoonomia_projection_dataset pipeline,
commit 4729d06d576f).
v4 re-derives the v3 partition with the labeling scheme resolved in
issue #221: base-pair
priority + window majority, a protein-coding-only TSS band, and
ccre_flank=0. See the pipeline README's "v4 region-type annotation"… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_bg.zoonomia-v1-v4_utr3
bolinas-dna/zoonomia-v1-v4_utr3
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled utr3 by the v4 region labeler
(snakemake/zoonomia_projection_dataset pipeline,
commit 4729d06d576f).
v4 re-derives the v3 partition with the labeling scheme resolved in
issue #221: base-pair
priority + window majority, a protein-coding-only TSS band, and
ccre_flank=0. See the pipeline README's "v4 region-type annotation"… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_utr3.zoonomia-v1-v2
zoonomia-v1-v2
Same projection backbone as bolinas-dna/zoonomia-v1-v1, subsetted to windows whose human anchor overlaps [TSS − 256, TSS + 256] for any Ensembl rel 115 protein_coding transcript. ~7.1% of v1.
Produced by the zoonomia_projection_dataset pipeline in Open-Athena/bolinas-dna — permalinked at the exact code that built this dataset: snakemake/zoonomia_projection_dataset @ 7ff07cd (PR #158).
Row count
Train: 15,863,610 rows.
Verified by zstd -dc… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v2.zoonomia-rag-v1-v1
bolinas-dna/zoonomia-rag-v1-v1
Fixed-layout 2,048-token documents built from conservation-filtered GRCh38
255-base anchors, seven fixed Zoonomia mammalian ortholog slots, and a final
human slot. Missing non-human projections are filled with 255 N bases;
chromosome 18 is validation-only.
Produced by the commit-pinned issue #402 RAG pipeline. The
immutable upstream input is the existing Zoonomia v1
min0.20/all_species_with_sequence.parquet projection. No halLiftover was
run for… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-rag-v1-v1.zoonomia-v1-v4_tss_region_and_utr5
bolinas-dna/zoonomia-v1-v4_tss_region_and_utr5
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled tss_region_and_utr5 by the v4 region labeler
(snakemake/zoonomia_projection_dataset pipeline,
commit 4729d06d576f).
v4 re-derives the v3 partition with the labeling scheme resolved in
issue #221: base-pair
priority + window majority, a protein-coding-only TSS band, and
ccre_flank=0. See the pipeline README's… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_tss_region_and_utr5.zoonomia-v1-v4_ncrna_exon-order
bolinas-dna/zoonomia-v1-v4_ncrna_exon-order
The bolinas-dna/zoonomia-v1-v4_ncrna_exon cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors and same… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ncrna_exon-order.zoonomia-v1-v3_bg
bolinas-dna/zoonomia-v1-v3_bg
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled background by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (background)
Background — anchors whose union-of-functional fraction over the five labels above is below 0.20 (or that have zero overlap with any). About 70% are intronic (gene-body but not exonic) and 30%… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_bg.zoonomia-v1-v4_utr3-order
bolinas-dna/zoonomia-v1-v4_utr3-order
The bolinas-dna/zoonomia-v1-v4_utr3 cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors and same per-window… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_utr3-order.zoonomia-v1-v4_ccre_enhancer_centered-order
bolinas-dna/zoonomia-v1-v4_ccre_enhancer_centered-order
An enhancer-CENTERED training set for issue
#351, built by the
snakemake/zoonomia_projection_dataset pipeline
(workflow/rules/centered.smk) at commit
8127acfea5aa.
Provenance
Each training window is defined directly from an ENCODE cCRE V4 enhancer
(dELS + pELS): one 255 bp window centered on the cCRE midpoint
(make_enhancer_anchors, keep-all — clustered enhancers each keep their own
window), rather than a… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_enhancer_centered-order.zoonomia-v1-v4_tss_region_and_utr5-order
bolinas-dna/zoonomia-v1-v4_tss_region_and_utr5-order
The bolinas-dna/zoonomia-v1-v4_tss_region_and_utr5 cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_tss_region_and_utr5-order.zoonomia-v1-val_ncrna
bolinas-dna/zoonomia-v1-val_ncrna
Conservation pre-filtered, case-encoded human-genome validation set from the
snakemake/zoonomia_projection_dataset pipeline
(commit main).
This is one of seven per-recipe validation parquets (val_cds, val_utr5,
val_utr3, val_ncrna, val_promoter, val_enhancer, val_tss_pc) built
from the same human-anchored phyloP_447m scoring used to create the
cross-mammal training sets
bolinas-dna/zoonomia-v1-v1
and bolinas-dna/zoonomia-v1-v2.… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-val_ncrna.zoonomia-v1-val_cds
bolinas-dna/zoonomia-v1-val_cds
Conservation pre-filtered, case-encoded human-genome validation set from the
snakemake/zoonomia_projection_dataset pipeline
(commit main).
This is one of seven per-recipe validation parquets (val_cds, val_utr5,
val_utr3, val_ncrna, val_promoter, val_enhancer, val_tss_pc) built
from the same human-anchored phyloP_447m scoring used to create the
cross-mammal training sets
bolinas-dna/zoonomia-v1-v1
and bolinas-dna/zoonomia-v1-v2.
Recipe… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-val_cds.zoonomia-v1-val_enhancer
bolinas-dna/zoonomia-v1-val_enhancer
Conservation pre-filtered, case-encoded human-genome validation set from the
snakemake/zoonomia_projection_dataset pipeline
(commit main).
This is one of seven per-recipe validation parquets (val_cds, val_utr5,
val_utr3, val_ncrna, val_promoter, val_enhancer, val_tss_pc) built
from the same human-anchored phyloP_447m scoring used to create the
cross-mammal training sets
bolinas-dna/zoonomia-v1-v1
and bolinas-dna/zoonomia-v1-v2.… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-val_enhancer.zoonomia-v1-val_utr3
bolinas-dna/zoonomia-v1-val_utr3
Conservation pre-filtered, case-encoded human-genome validation set from the
snakemake/zoonomia_projection_dataset pipeline
(commit main).
This is one of seven per-recipe validation parquets (val_cds, val_utr5,
val_utr3, val_ncrna, val_promoter, val_enhancer, val_tss_pc) built
from the same human-anchored phyloP_447m scoring used to create the
cross-mammal training sets
bolinas-dna/zoonomia-v1-v1
and bolinas-dna/zoonomia-v1-v2.
Recipe… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-val_utr3.zoonomia-v1-val_utr5
bolinas-dna/zoonomia-v1-val_utr5
Conservation pre-filtered, case-encoded human-genome validation set from the
snakemake/zoonomia_projection_dataset pipeline
(commit main).
This is one of seven per-recipe validation parquets (val_cds, val_utr5,
val_utr3, val_ncrna, val_promoter, val_enhancer, val_tss_pc) built
from the same human-anchored phyloP_447m scoring used to create the
cross-mammal training sets
bolinas-dna/zoonomia-v1-v1
and bolinas-dna/zoonomia-v1-v2.
Recipe… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-val_utr5.zoonomia-v1-val_promoter
bolinas-dna/zoonomia-v1-val_promoter
Conservation pre-filtered, case-encoded human-genome validation set from the
snakemake/zoonomia_projection_dataset pipeline
(commit main).
This is one of seven per-recipe validation parquets (val_cds, val_utr5,
val_utr3, val_ncrna, val_promoter, val_enhancer, val_tss_pc) built
from the same human-anchored phyloP_447m scoring used to create the
cross-mammal training sets
bolinas-dna/zoonomia-v1-v1
and bolinas-dna/zoonomia-v1-v2.… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-val_promoter.
