datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zoonomia-v1-v3_ncrna_exon
bolinas-dna/zoonomia-v1-v3_ncrna_exon
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ncrna_exon by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (ncrna_exon)
Non-coding-RNA exon — every Ensembl r115 exon that is not part of a protein-coding transcript (get_exons(ann) − get_ensembl_protein_coding_exons(ann)). No biotype or quality filter, so this… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ncrna_exon.phylop-uniform-v1-ncrna-exon
marin-dna/phylop-uniform-v1-ncrna-exon
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the ncrna_exon region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-ncrna-exon.gpn-star-p-uniform-v1-ncrna-exon
marin-dna/gpn-star-p-uniform-v1-ncrna-exon
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the ncrna_exon region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-ncrna-exon.zoonomia-v1-v4_ncrna_exon
bolinas-dna/zoonomia-v1-v4_ncrna_exon
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ncrna_exon by the v4 region labeler
(snakemake/zoonomia_projection_dataset pipeline,
commit 4729d06d576f).
v4 re-derives the v3 partition with the labeling scheme resolved in
issue #221: base-pair
priority + window majority, a protein-coding-only TSS band, and
ccre_flank=0. See the pipeline README's "v4 region-type… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ncrna_exon.functional-ncrna
marin-dna/functional-ncrna
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains.
This draft covers the ncrna region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
For the 28 non-mammalian targets, the stable… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-ncrna.zoonomia-v1-v4_ncrna_exon-order
bolinas-dna/zoonomia-v1-v4_ncrna_exon-order
The bolinas-dna/zoonomia-v1-v4_ncrna_exon cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors and same… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ncrna_exon-order.vertebrate-v1-ncrna_exon
marin-dna/vertebrate-v1-ncrna_exon
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the ncrna_exon region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-ncrna_exon.
