datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genomes-v4-genome_set-animals-intervals-v5_256_128genomes-v4-genome_set-animals-intervals-v11_256_128genomes-v4-genome_set-animals-intervals-v4_512_256genomes-v5-genome_set-animals-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128
Animals promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
68,286,166 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.genomes-v4-genome_set-animals-intervals-v12_256_128genomes-v4-genome_set-animals-intervals-v10_256_128genomes-v4-genome_set-animals-intervals-v13_256_128genomes-v4-genome_set-animals-intervals-v14_256_128genomes-v5-genome_set-animals_order204-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128
204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
main). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
101,114,252 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v1_256_128genomes-v4-genome_set-animals-intervals-v6_256_128genomes-v5-genome_set-animals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128
Animals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
242,334,716 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v8_256_128genomes-v4-genome_set-animals-intervals-v9_256_128genomes-v4-genome_set-animals-intervals-v15_256_128genomes-v5-genome_set-animals-intervals-v15_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v15_255_128
Animals downstream-of-CDS (v15) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
20,501,856 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v15_255_128.genomes-v3-genome_set-animals-intervals-v3_512_256genomes-v3-genome_set-animals-intervals-v2_512_256genomes-v2-genome_set-animals-intervals-v1_512_256animals-cds-proj-v1-vert125
Animal CDS via human→animal projection — vertebrates (Chordata)
Human protein-coding (CDS) 255 bp windows projected onto 125 vertebrate species (Chordata, one per order) by mmseqs2 nucleotide local alignment — the projection ("Arm B") side of the issue #353 CDS projection-vs-annotation experiment.
How it was built
The human v5 CDS intervals (Ensembl coding exons, filtered 20–10,000 bp, +20 bp splice flank, expanded ≥256 bp) are tiled into 255 bp windows (269,866… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/animals-cds-proj-v1-vert125.animals-cds-proj-v1-all204
Animal CDS via human→animal projection — all animal orders
Human protein-coding (CDS) 255 bp windows projected onto 204 animal species (one annotated genome per Metazoan order) by mmseqs2 nucleotide local alignment — the projection ("Arm B") side of the issue #353 CDS projection-vs-annotation experiment.
How it was built
The human v5 CDS intervals (Ensembl coding exons, filtered 20–10,000 bp, +20 bp splice flank, expanded ≥256 bp) are tiled into 255 bp windows… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/animals-cds-proj-v1-all204.genomes-v4-genome_set-animals-intervals-v1_512_256ai-4-animals-v0AnimalSnaps-Labeled
AnimalSnaps-Labeled
Collection Protocol
Images were collected by volunteers using mobile phones in public parks between March and May 2025. Each image was independently labeled by two annotators; disagreements were resolved by a senior annotator. Blurry or duplicate images were discarded before labeling.
Label Distribution
Label
Count
Percentage
cat
30
37.5%
dog
25
31.2%
bird
15
18.8%
fox
10
12.5%
AnimalSnaps-Labeled
AnimalSnaps-Labeled
Collection Protocol
Images were collected by volunteers using mobile phones in public parks between March and May 2025. Each image was independently labeled by two annotators; disagreements were resolved by a senior annotator. Blurry or duplicate images were discarded before labeling.
Label Distribution
Label
Count
Percentage
cat
30
37.5%
dog
25
31.2%
bird
15
18.8%
fox
10
12.5%
genomes-v2-genome_set-animals-intervals-v2_512_256Synthetic-Animalssubliminal-animals-smoketestindonesia_animals_datasetanimals-80-QA
