CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-dna /genomes-v4-genome_set-animals-intervals-v5_256_128text100M<n<1B0 likes1.5k downloads8mo agoHugging Face02marin-dna /genomes-v4-genome_set-animals-intervals-v11_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face03marin-dna /genomes-v4-genome_set-animals-intervals-v10_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face04marin-dna /genomes-v4-genome_set-animals-intervals-v12_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face05marin-dna /genomes-v5-genome_set-animals_order204-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128 204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit main). Each repo in the family is one (genome_set, region-recipe) combination. Size 101,114,252 sequences across 64 data/train/*.jsonl.zst shards (reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.text100M<n<1B0 likes1.3k downloads3mo agoHugging Face06marin-dna /genomes-v4-genome_set-animals-intervals-v13_256_128text10M<n<100M0 likes1.3k downloads8mo agoHugging Face07marin-dna /genomes-v4-genome_set-animals-intervals-v4_512_256text10M<n<100M0 likes1.3k downloads8mo agoHugging Face08marin-dna /genomes-v4-genome_set-animals-intervals-v14_256_128text10M<n<100M0 likes1.3k downloads8mo agoHugging Face09marin-dna /genomes-v5-genome_set-animals-intervals-v1_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128 Animals promoters (v1) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 68,286,166 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.text10M<n<100M0 likes1.3k downloads29d agoHugging Face10marin-dna /genomes-v4-genome_set-animals-intervals-v1_256_128text10M<n<100M0 likes1.2k downloads8mo agoHugging Face11marin-dna /zoonomia-v1-v4_ccre_noexon bolinas-dna/zoonomia-v1-v4_ccre_noexon A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is the v4 ccre_non_promoter arm with every window that overlaps any other functional element (CDS / 3′UTR / ncRNA exon / TSS+5′UTR) removed — i.e. windows whose functional content… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon.tabular10M<n<100M0 likes1.2k downloads3mo agoHugging Face12marin-dna /zoonomia-v1-v4_ccre_noexon_enhancer bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.tabular10M<n<100M0 likes1.1k downloads3mo agoHugging Face13marin-dna /zoonomia-v1-v3_cds bolinas-dna/zoonomia-v1-v3_cds Per-anchor region-type partition of the cross-mammal training set bolinas-dna/zoonomia-v1-v1, restricted to anchors labelled cds by the snakemake/zoonomia_projection_dataset pipeline (commit 2ab868a2f1d4). Region label (cds) Coding sequence — Ensembl r115 CDS features (get_cds). Highest-priority class: any anchor with overlap on a CDS feature (and union-of-functional fraction ≥ 0.20 across all five labels) is labelled cds, regardless… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_cds.tabular10M<n<100M0 likes1.1k downloads5mo agoHugging Face14marin-dna /genomes-v4-genome_set-animals-intervals-v6_256_128text10M<n<100M0 likes1.1k downloads8mo agoHugging Face15marin-dna /gpn-star-p-uniform-v1-cds marin-dna/gpn-star-p-uniform-v1-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-cds.tabular10M<n<100M0 likes1k downloads1mo agoHugging Face16marin-dna /phylop-uniform-v1-cds marin-dna/phylop-uniform-v1-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-cds.tabular10M<n<100M0 likes1k downloads29d agoHugging Face17marin-dna /genomes-v5-genome_set-animals-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128 Animals CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 242,334,716 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.text100M<n<1B0 likes1k downloads29d agoHugging Face18marin-dna /gpn-star-p-uniform-v1-enhancer-arm-a marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.tabular100M<n<1B0 likes966 downloads1mo agoHugging Face19marin-dna /functional-cds marin-dna/functional-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. For the 28 non-mammalian targets, the stable ucsc_multiz100way… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-cds.tabular10M<n<100M0 likes938 downloads1mo agoHugging Face20marin-dna /genomes-v4-genome_set-animals-intervals-v8_256_128text10M<n<100M0 likes928 downloads8mo agoHugging Face21marin-dna /zoonomia-v1-v1 zoonomia-v1-v1 255 bp human-anchored windows, conservation-filtered (phyloP_447m, proportion_conserved >= 0.20), projected onto 108 family-deduped Zoonomia 447-mammalian assemblies via halLiftover, midpoint-resized to 255 bp, reverse-complement-augmented. Single train split. Produced by the zoonomia_projection_dataset pipeline in Open-Athena/bolinas-dna — permalinked at the exact code that built this dataset: snakemake/zoonomia_projection_dataset @ 7ff07cd (PR #158).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v1.tabular100M<n<1B0 likes904 downloads5mo agoHugging Face22marin-dna /genomes-v4-genome_set-animals-intervals-v9_256_128text10M<n<100M0 likes888 downloads8mo agoHugging Face23marin-dna /genomes-v4-genome_set-animals-intervals-v15_256_128text10M<n<100M0 likes887 downloads8mo agoHugging Face24marin-dna /zoonomia-v1-v3_ccre_non_promoter bolinas-dna/zoonomia-v1-v3_ccre_non_promoter Per-anchor region-type partition of the cross-mammal training set bolinas-dna/zoonomia-v1-v1, restricted to anchors labelled ccre_non_promoter by the snakemake/zoonomia_projection_dataset pipeline (commit 2ab868a2f1d4). Region label (ccre_non_promoter) ENCODE cCRE V4 non-promoter classes — cre_class != "PLS" (so: dELS, pELS, CA, CA-CTCF, CA-TF, CA-H3K4me3, TF), extended by 500 bp on each side. PLS is excluded because… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ccre_non_promoter.tabular10M<n<100M0 likes886 downloads29d agoHugging Face25marin-dna /vertebrate-v1-issue473-fullwindow-cds-random-val marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val CDS full-window vertebrate projection sequences for the issue #473 random validation control. The source is the immutable issue #417 accepted-sequence table. The split uniformly samples 16,384 original-orientation CDS rows without replacement using seed 42. Sampling occurs before reverse-complement augmentation. Selected rows are removed from training; reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.tabular10M<n<100M0 likes883 downloads1mo agoHugging Face26marin-community /mcp-atlas-easy MCP-Atlas-Easy An easy, single-tool-call benchmark for pretrained (base) language models, derived from ScaleAI/MCP-Atlas. MCP-Atlas evaluates instruction-tuned agents on multi-step tool orchestration (3–6 calls per task across 36 real MCP servers). MCP-Atlas-Easy strips that down to the simplest possible form of the same skill: one tool spec, one trivially unambiguous request, one correct tool call, then stop. This makes it usable as a completion-style eval for base models with… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/mcp-atlas-easy.texttext-generationn<1K1 likes796 downloads2mo agoHugging Face27marin-dna /gpn-star-p-uniform-v1-background marin-dna/gpn-star-p-uniform-v1-background Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-background.tabular10M<n<100M0 likes786 downloads1mo agoHugging Face28marin-dna /genomes-v5-genome_set-vertebrates-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128 Vertebrates CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 172,025,122 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128.text100M<n<1B0 likes770 downloads4mo agoHugging Face29marin-dna /vertebrate-v1-all marin-dna/vertebrate-v1-all Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the all region cohort with all species scope and preserves source FASTA/2bit letter case. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence case is independent of that filter: lowercase bases preserve source repeat masking, uppercase bases preserve source non-repeat-masked… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-all.tabular100M<n<1B0 likes761 downloads2mo agoHugging Face30marin-dna /phylop-uniform-v1-enhancer-arm-a marin-dna/phylop-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.tabular10M<n<100M0 likes761 downloads29d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.