CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-community /fineweb-edu-pretokenized-10K Marin/Levanter Subsampled Pretokenized Dataset Dataset Train Urls: gs://marin-us-central2/raw/fineweb-edu-c2beb4/3c452cb/huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/3c452cb Factsheet Original cache: gs://marin-us-central2/tokenized/fineweb-edu-24698d Tokenizer: stanford-crfm/marin-tokenizer Seed 42 Number of tokens: 10,390 (This readme is automatically generated by Marin.) 0 likes4k downloads1y agoHugging Face02open-athena /marinskyrl-gpu-wheelhouse0 likes2.6k downloads9d agoHugging Face03marin-community /swe-rebench-v2-CodeWorldModeling SWE-rebench V2 — CodeWorldModeling Traces This is a derived dataset. Every record is produced from an instance of nebius/SWE-rebench-V2. It is governed by the SWE-rebench V2 license — see License below — including the requirement to respect each source repository's own license. Line-by-line Python execution traces for the test suites of SWE-rebench V2 instances, captured by running each instance's tests under a tracer inside Nebius ConTree sandboxes. Each instance comes with a fix… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/swe-rebench-v2-CodeWorldModeling.texttext-generation1M<n<10M2 likes1.7k downloads4mo agoHugging Face04marin-dna /genomes-v4-genome_set-animals-intervals-v5_256_128text100M<n<1B0 likes1.5k downloads8mo agoHugging Face05AnTo2209 /MarineEVT MarineEVT Dataset MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning 📖 Description MarineEVT is a comprehensive event-centric dataset and benchmark for marine video understanding. It comprises 20,000 richly annotated underwater video question-answer pairs spanning 20 fine-grained dimensions, designed to support semantic, contextualized, spatial-temporal, and causal reasoning in marine environments. The dataset addresses the… See the full description on the dataset page: https://huggingface.co/datasets/AnTo2209/MarineEVT.0 likes1.4k downloads3mo agoHugging Face06marin-dna /genomes-v4-genome_set-animals-intervals-v11_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face07marin-dna /genomes-v4-genome_set-animals-intervals-v10_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face08marin-community /grug-moe-mix-swarm Grug-MoE Data-Mix Experiments The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below. Fisher-DSP swarm (default) 840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2). Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.tabulartext-generation1K<n<10K1 likes1.3k downloads13d agoHugging Face09marin-dna /genomes-v4-genome_set-animals-intervals-v12_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face10marin-dna /genomes-v5-genome_set-animals_order204-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128 204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit main). Each repo in the family is one (genome_set, region-recipe) combination. Size 101,114,252 sequences across 64 data/train/*.jsonl.zst shards (reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.text100M<n<1B0 likes1.3k downloads3mo agoHugging Face11marin-dna /genomes-v4-genome_set-animals-intervals-v13_256_128text10M<n<100M0 likes1.3k downloads8mo agoHugging Face12marin-dna /genomes-v4-genome_set-animals-intervals-v4_512_256text10M<n<100M0 likes1.3k downloads8mo agoHugging Face13marin-community /token-counts Marin Token Counts Token counts for all datasets used in Marin pretraining runs. Schema Column Type Description dataset string Dataset identifier marin_tokens int Number of tokens after tokenization category string Content domain (web, code, math, academic, books, etc.) synthetic bool Whether the data is LLM-generated or LLM-translated Categories web — Quality-classified Common Crawl text (Nemotron-CC) code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.texttext-generationn<1K1 likes1.3k downloads29d agoHugging Face14marin-dna /genomes-v4-genome_set-animals-intervals-v14_256_128text10M<n<100M0 likes1.3k downloads8mo agoHugging Face15marin-dna /genomes-v4-genome_set-animals-intervals-v7_256_1280 likes1.3k downloads8mo agoHugging Face16marin-dna /genomes-v5-genome_set-animals-intervals-v1_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128 Animals promoters (v1) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 68,286,166 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.text10M<n<100M0 likes1.3k downloads29d agoHugging Face17marin-dna /genomes-v4-genome_set-animals-intervals-v1_256_128text10M<n<100M0 likes1.2k downloads8mo agoHugging Face18marin-dna /zoonomia-v1-v4_ccre_noexon bolinas-dna/zoonomia-v1-v4_ccre_noexon A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is the v4 ccre_non_promoter arm with every window that overlaps any other functional element (CDS / 3′UTR / ncRNA exon / TSS+5′UTR) removed — i.e. windows whose functional content… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon.tabular10M<n<100M0 likes1.2k downloads3mo agoHugging Face19DeepSeekOracle /lattice-marines-wins Lattice Marines Eternal Ledger Public vs-AI win book for Lattice Marines. Hall: https://chatagent.ca/games/lattice-marines/ledger.html Writer Space: https://huggingface.co/spaces/DeepSeekOracle/lattice-marines-ledger File: ledger.json Only adaptive-AI victories are stored (commander name, score, difficulty, map size, seed, turns, AI profile, date). Hot-seat is excluded. 0 likes1.2k downloads12h agoHugging Face20marin-dna /zoonomia-v1-v4_ccre_noexon_enhancer bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.tabular10M<n<100M0 likes1.1k downloads3mo agoHugging Face21marin-dna /zoonomia-v1-v3_cds bolinas-dna/zoonomia-v1-v3_cds Per-anchor region-type partition of the cross-mammal training set bolinas-dna/zoonomia-v1-v1, restricted to anchors labelled cds by the snakemake/zoonomia_projection_dataset pipeline (commit 2ab868a2f1d4). Region label (cds) Coding sequence — Ensembl r115 CDS features (get_cds). Highest-priority class: any anchor with overlap on a CDS feature (and union-of-functional fraction ≥ 0.20 across all five labels) is labelled cds, regardless… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_cds.tabular10M<n<100M0 likes1.1k downloads5mo agoHugging Face22marin-dna /genomes-v4-genome_set-animals-intervals-v6_256_128text10M<n<100M0 likes1.1k downloads8mo agoHugging Face23marin-dna /gpn-star-p-uniform-v1-cds marin-dna/gpn-star-p-uniform-v1-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-cds.tabular10M<n<100M0 likes1k downloads1mo agoHugging Face24marin-dna /phylop-uniform-v1-cds marin-dna/phylop-uniform-v1-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-cds.tabular10M<n<100M0 likes1k downloads29d agoHugging Face25marin-dna /genomes-v5-genome_set-animals-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128 Animals CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 242,334,716 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.text100M<n<1B0 likes1k downloads29d agoHugging Face26marin-dna /gpn-star-p-uniform-v1-enhancer-arm-a marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.tabular100M<n<1B0 likes966 downloads1mo agoHugging Face27eczech /marinfold-exp11-protein-docs-seq marinfold-exp11-pdocs-seq Sequence-only derivative of eczech/marinfold-exp11-protein-docs. For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.tabulartext-generation1M<n<10M0 likes961 downloads4mo agoHugging Face28marin-dna /functional-cds marin-dna/functional-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. For the 28 non-mammalian targets, the stable ucsc_multiz100way… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-cds.tabular10M<n<100M0 likes938 downloads1mo agoHugging Face29marin-dna /genomes-v4-genome_set-animals-intervals-v8_256_128text10M<n<100M0 likes928 downloads8mo agoHugging Face30Hariprasath5128 /marine-animals-multimodal-dataset Marine Animals Multimodal Dataset 🐋 A comprehensive multimodal dataset combining audio recordings and images of 32 marine species. Dataset Summary Total samples: 24,911 Species: 32 Audio files: 1,357 unique recordings Images: 581 (309 matched + 272 from iNaturalist) Features species (string): Species name label (int32): Numeric label (0–31) audio (Audio): Audio recording of the species image (Image): Species image image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.audioaudio-classification10K<n<100K0 likes904 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.