datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grug-moe-mix-swarm
Grug-MoE Data-Mix Experiments
The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below.
Fisher-DSP swarm (default)
840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2).
Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress
mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.zoonomia-v1-v4_ccre_noexon
bolinas-dna/zoonomia-v1-v4_ccre_noexon
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is the v4 ccre_non_promoter arm with every window that overlaps any other functional element (CDS / 3′UTR / ncRNA exon / TSS+5′UTR) removed — i.e. windows whose functional content… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon.zoonomia-v1-v4_ccre_noexon_enhancer
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.zoonomia-v1-v3_cds
bolinas-dna/zoonomia-v1-v3_cds
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled cds by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (cds)
Coding sequence — Ensembl r115 CDS features (get_cds). Highest-priority class: any anchor with overlap on a CDS feature (and union-of-functional fraction ≥ 0.20 across all five labels) is labelled cds, regardless… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_cds.gpn-star-p-uniform-v1-cds
marin-dna/gpn-star-p-uniform-v1-cds
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-cds.phylop-uniform-v1-cds
marin-dna/phylop-uniform-v1-cds
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-cds.gpn-star-p-uniform-v1-enhancer-arm-a
marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.marinfold-exp11-protein-docs-seq
marinfold-exp11-pdocs-seq
Sequence-only derivative of
eczech/marinfold-exp11-protein-docs.
For every row, the document field has been reduced to just the amino-acid sequence
portion: the <begin_sequence> tag followed by the per-residue three-letter tokens
(e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1>
document-type prefix and everything from <begin_statements> onward (contacts and
distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.functional-cds
marin-dna/functional-cds
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains.
This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
For the 28 non-mammalian targets, the stable ucsc_multiz100way… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-cds.zoonomia-v1-v1
zoonomia-v1-v1
255 bp human-anchored windows, conservation-filtered (phyloP_447m, proportion_conserved >= 0.20), projected onto 108 family-deduped Zoonomia 447-mammalian assemblies via halLiftover, midpoint-resized to 255 bp, reverse-complement-augmented. Single train split.
Produced by the zoonomia_projection_dataset pipeline in Open-Athena/bolinas-dna — permalinked at the exact code that built this dataset: snakemake/zoonomia_projection_dataset @ 7ff07cd (PR #158).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v1.zoonomia-v1-v3_ccre_non_promoter
bolinas-dna/zoonomia-v1-v3_ccre_non_promoter
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ccre_non_promoter by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (ccre_non_promoter)
ENCODE cCRE V4 non-promoter classes — cre_class != "PLS" (so: dELS, pELS, CA, CA-CTCF, CA-TF, CA-H3K4me3, TF), extended by 500 bp on each side. PLS is excluded because… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ccre_non_promoter.vertebrate-v1-issue473-fullwindow-cds-random-val
marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val
CDS full-window vertebrate projection sequences for the issue #473 random
validation control. The source is the immutable issue #417 accepted-sequence
table.
The split uniformly samples 16,384 original-orientation CDS rows
without replacement using seed 42. Sampling occurs before
reverse-complement augmentation. Selected rows are removed from training;
reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.gpn-star-p-uniform-v1-background
marin-dna/gpn-star-p-uniform-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-background.vertebrate-v1-all
marin-dna/vertebrate-v1-all
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the all region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source non-repeat-masked… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-all.phylop-uniform-v1-enhancer-arm-a
marin-dna/phylop-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.vertebrate-v1-issue473-center1-cds
marin-dna/vertebrate-v1-issue473-center1-cds
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
cds cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed #417 protein-coding CDS anchor catalog.
The source projection was produced by the… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-cds.functional-enhancer
marin-dna/functional-enhancer
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
For the 28 non-mammalian targets, the stable… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-enhancer.phylop-uniform-v1-ncrna-exon
marin-dna/phylop-uniform-v1-ncrna-exon
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the ncrna_exon region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-ncrna-exon.zoonomia-v1-v3_ncrna_exon
bolinas-dna/zoonomia-v1-v3_ncrna_exon
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ncrna_exon by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (ncrna_exon)
Non-coding-RNA exon — every Ensembl r115 exon that is not part of a protein-coding transcript (get_exons(ann) − get_ensembl_protein_coding_exons(ann)). No biotype or quality filter, so this… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ncrna_exon.gpn-star-p-uniform-v1-utr3
marin-dna/gpn-star-p-uniform-v1-utr3
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the utr3 region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-utr3.open-thoughts-4-math-qwen3-32b-annotated
Dataset Card for Open-Thoughts-4-Math-Qwen3-32B-Annotated
This dataset is the Qwen3-32B annotated version of mlfoundations-dev/hero_run_4_math curated by the
OpenThoughts4 team. We provide the responses from Qwen3-32B in the generated_text column. These samples were generated using temperature = 0.8 and max output tokens = 7,500.
We note that many of the responses are truncated, so use this dataset wisely!
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-math-qwen3-32b-annotated.marinfold-exp11-protein-docs
marinfold-exp11-pdocs
Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from
timodonnell/protein-docs,
partitioned by the source round column:
Config
Source rounds
Approx rows
high
round 0
~1.68M
medium
round 1
~1.42M
low
round 2–4
~2.29M
Train/val/test split assignment is inherited from the source dataset (leakage-resistant
structural-cluster hashing). All columns from the source are preserved; rows are simply
partitioned by round.
See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.zoonomia-v1-v4_ccre_non_promoter-order
bolinas-dna/zoonomia-v1-v4_ccre_non_promoter-order
The bolinas-dna/zoonomia-v1-v4_ccre_non_promoter cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_non_promoter-order.gpn-star-p-uniform-v1-tss-utr5
marin-dna/gpn-star-p-uniform-v1-tss-utr5
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the tss_region_and_utr5 region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-tss-utr5.vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the full_window policy.
The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.vertebrate-v1-cds_mammals_only
marin-dna/vertebrate-v1-cds_mammals_only
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment. This
draft covers the cds region cohort with mammals_only species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source non-repeat-masked sequence, and… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-cds_mammals_only.zoonomia-v1-v4_ccre_non_promoter
bolinas-dna/zoonomia-v1-v4_ccre_non_promoter
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ccre_non_promoter by the v4 region labeler
(snakemake/zoonomia_projection_dataset pipeline,
commit 4729d06d576f).
v4 re-derives the v3 partition with the labeling scheme resolved in
issue #221: base-pair
priority + window majority, a protein-coding-only TSS band, and
ccre_flank=0. See the pipeline README's "v4… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_non_promoter.openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.vertebrate-v1-ccre_non_promoter
marin-dna/vertebrate-v1-ccre_non_promoter
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the ccre_non_promoter region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-ccre_non_promoter.openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-32B (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-32B
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16.
