CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-dna /genomes-v4-genome_set-animals-intervals-v5_256_128text100M<n<1B0 likes1.2k downloads8mo agoHugging Face02marin-dna /genomes-v4-genome_set-animals-intervals-v4_512_256text10M<n<100M0 likes1.1k downloads8mo agoHugging Face03marin-dna /genomes-v4-genome_set-animals-intervals-v11_256_128text100M<n<1B0 likes1k downloads8mo agoHugging Face04marin-dna /genomes-v5-genome_set-animals-intervals-v1_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128 Animals promoters (v1) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 68,286,166 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.text10M<n<100M0 likes1k downloads26d agoHugging Face05marin-dna /genomes-v4-genome_set-animals-intervals-v10_256_128text100M<n<1B0 likes1k downloads8mo agoHugging Face06marin-dna /genomes-v4-genome_set-animals-intervals-v12_256_128text100M<n<1B0 likes1k downloads8mo agoHugging Face07marin-dna /genomes-v4-genome_set-animals-intervals-v13_256_128text10M<n<100M0 likes1k downloads8mo agoHugging Face08marin-dna /genomes-v4-genome_set-animals-intervals-v14_256_128text10M<n<100M0 likes1k downloads8mo agoHugging Face09marin-dna /genomes-v5-genome_set-animals_order204-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128 204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit main). Each repo in the family is one (genome_set, region-recipe) combination. Size 101,114,252 sequences across 64 data/train/*.jsonl.zst shards (reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.text100M<n<1B0 likes946 downloads3mo agoHugging Face10marin-dna /genomes-v4-genome_set-animals-intervals-v1_256_128text10M<n<100M0 likes944 downloads8mo agoHugging Face11marin-dna /gpn-star-p-uniform-v1-cds marin-dna/gpn-star-p-uniform-v1-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-cds.tabular10M<n<100M0 likes903 downloads28d agoHugging Face12marin-dna /genomes-v4-genome_set-animals-intervals-v6_256_128text10M<n<100M0 likes873 downloads8mo agoHugging Face13marin-dna /zoonomia-v1-v4_ccre_noexon bolinas-dna/zoonomia-v1-v4_ccre_noexon A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is the v4 ccre_non_promoter arm with every window that overlaps any other functional element (CDS / 3′UTR / ncRNA exon / TSS+5′UTR) removed — i.e. windows whose functional content… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon.tabular10M<n<100M0 likes863 downloads3mo agoHugging Face14marin-dna /gpn-star-p-uniform-v1-enhancer-arm-a marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.tabular100M<n<1B0 likes856 downloads28d agoHugging Face15marin-dna /gpn-star-p-uniform-v1-background marin-dna/gpn-star-p-uniform-v1-background Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-background.tabular10M<n<100M0 likes855 downloads28d agoHugging Face16marin-dna /phylop-uniform-v1-cds marin-dna/phylop-uniform-v1-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-cds.tabular10M<n<100M0 likes840 downloads26d agoHugging Face17marin-dna /zoonomia-v1-v4_ccre_noexon_enhancer bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.tabular10M<n<100M0 likes830 downloads3mo agoHugging Face18marin-dna /genomes-v4-genome_set-animals-intervals-v8_256_128text10M<n<100M0 likes828 downloads8mo agoHugging Face19marin-dna /functional-cds marin-dna/functional-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. For the 28 non-mammalian targets, the stable ucsc_multiz100way… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-cds.tabular10M<n<100M0 likes816 downloads29d agoHugging Face20marin-dna /genomes-v5-genome_set-animals-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128 Animals CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 242,334,716 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.text100M<n<1B0 likes814 downloads26d agoHugging Face21marin-dna /zoonomia-v1-v3_cds bolinas-dna/zoonomia-v1-v3_cds Per-anchor region-type partition of the cross-mammal training set bolinas-dna/zoonomia-v1-v1, restricted to anchors labelled cds by the snakemake/zoonomia_projection_dataset pipeline (commit 2ab868a2f1d4). Region label (cds) Coding sequence — Ensembl r115 CDS features (get_cds). Highest-priority class: any anchor with overlap on a CDS feature (and union-of-functional fraction ≥ 0.20 across all five labels) is labelled cds, regardless… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_cds.tabular10M<n<100M0 likes810 downloads4mo agoHugging Face22marin-dna /genomes-v4-genome_set-animals-intervals-v9_256_128text10M<n<100M0 likes805 downloads8mo agoHugging Face23marin-dna /genomes-v4-genome_set-animals-intervals-v15_256_128text10M<n<100M0 likes783 downloads8mo agoHugging Face24marin-dna /genomes-v5-genome_set-animals-intervals-v15_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v15_255_128 Animals downstream-of-CDS (v15) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 20,501,856 sequences across 64 data/train/*.jsonl.zst shards (reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v15_255_128.text10M<n<100M0 likes721 downloads26d agoHugging Face25marin-dna /vertebrate-v1-issue473-fullwindow-cds-random-val marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val CDS full-window vertebrate projection sequences for the issue #473 random validation control. The source is the immutable issue #417 accepted-sequence table. The split uniformly samples 16,384 original-orientation CDS rows without replacement using seed 42. Sampling occurs before reverse-complement augmentation. Selected rows are removed from training; reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.tabular10M<n<100M0 likes712 downloads1mo agoHugging Face26marin-dna /genomes-v4-genome_set-mammals-intervals-v1_255_128-id1_cov1text10M<n<100M0 likes675 downloads7mo agoHugging Face27marin-dna /zoonomia-v1-v3_ccre_non_promoter bolinas-dna/zoonomia-v1-v3_ccre_non_promoter Per-anchor region-type partition of the cross-mammal training set bolinas-dna/zoonomia-v1-v1, restricted to anchors labelled ccre_non_promoter by the snakemake/zoonomia_projection_dataset pipeline (commit 2ab868a2f1d4). Region label (ccre_non_promoter) ENCODE cCRE V4 non-promoter classes — cre_class != "PLS" (so: dELS, pELS, CA, CA-CTCF, CA-TF, CA-H3K4me3, TF), extended by 500 bp on each side. PLS is excluded because… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ccre_non_promoter.tabular10M<n<100M0 likes671 downloads26d agoHugging Face28marin-dna /functional-enhancer marin-dna/functional-enhancer Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. For the 28 non-mammalian targets, the stable… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-enhancer.tabular10M<n<100M0 likes649 downloads29d agoHugging Face29marin-dna /zoonomia-v1-v1 zoonomia-v1-v1 255 bp human-anchored windows, conservation-filtered (phyloP_447m, proportion_conserved >= 0.20), projected onto 108 family-deduped Zoonomia 447-mammalian assemblies via halLiftover, midpoint-resized to 255 bp, reverse-complement-augmented. Single train split. Produced by the zoonomia_projection_dataset pipeline in Open-Athena/bolinas-dna — permalinked at the exact code that built this dataset: snakemake/zoonomia_projection_dataset @ 7ff07cd (PR #158).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v1.tabular100M<n<1B0 likes618 downloads5mo agoHugging Face30marin-dna /zoonomia-v1-v3_ncrna_exon bolinas-dna/zoonomia-v1-v3_ncrna_exon Per-anchor region-type partition of the cross-mammal training set bolinas-dna/zoonomia-v1-v1, restricted to anchors labelled ncrna_exon by the snakemake/zoonomia_projection_dataset pipeline (commit 2ab868a2f1d4). Region label (ncrna_exon) Non-coding-RNA exon — every Ensembl r115 exon that is not part of a protein-coding transcript (get_exons(ann) − get_ensembl_protein_coding_exons(ann)). No biotype or quality filter, so this… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ncrna_exon.tabular10M<n<100M0 likes613 downloads26d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.