datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genomes-v4-genome_set-animals-intervals-v5_256_128genomes-v4-genome_set-animals-intervals-v11_256_128genomes-v4-genome_set-animals-intervals-v10_256_128genomes-v4-genome_set-animals-intervals-v12_256_128genomes-v5-genome_set-animals_order204-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128
204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
main). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
101,114,252 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v13_256_128genomes-v4-genome_set-animals-intervals-v4_512_256genomes-v4-genome_set-animals-intervals-v14_256_128genomes-v5-genome_set-animals-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128
Animals promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
68,286,166 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.genomes-v4-genome_set-animals-intervals-v1_256_128genomes-v4-genome_set-animals-intervals-v6_256_128genomes-v5-genome_set-animals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128
Animals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
242,334,716 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.gpn-animal-promoter-datasetgenomes-v4-genome_set-animals-intervals-v8_256_128genomes-v4-genome_set-animals-intervals-v9_256_128genomes-v4-genome_set-animals-intervals-v15_256_128genomes-v5-genome_set-animals-intervals-v15_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v15_255_128
Animals downstream-of-CDS (v15) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
20,501,856 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v15_255_128.genomes-v3-genome_set-animals-intervals-v2_512_256genomes-v3-genome_set-animals-intervals-v3_512_256genomes-v2-genome_set-animals-intervals-v1_512_256animals-cds-proj-v1-vert125
Animal CDS via human→animal projection — vertebrates (Chordata)
Human protein-coding (CDS) 255 bp windows projected onto 125 vertebrate species (Chordata, one per order) by mmseqs2 nucleotide local alignment — the projection ("Arm B") side of the issue #353 CDS projection-vs-annotation experiment.
How it was built
The human v5 CDS intervals (Ensembl coding exons, filtered 20–10,000 bp, +20 bp splice flank, expanded ≥256 bp) are tiled into 255 bp windows (269,866… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/animals-cds-proj-v1-vert125.animals-cds-proj-v1-all204
Animal CDS via human→animal projection — all animal orders
Human protein-coding (CDS) 255 bp windows projected onto 204 animal species (one annotated genome per Metazoan order) by mmseqs2 nucleotide local alignment — the projection ("Arm B") side of the issue #353 CDS projection-vs-annotation experiment.
How it was built
The human v5 CDS intervals (Ensembl coding exons, filtered 20–10,000 bp, +20 bp splice flank, expanded ≥256 bp) are tiled into 255 bp windows… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/animals-cds-proj-v1-all204.genomes-v4-genome_set-animals-intervals-v1_512_256animal-welfare-training-claude
Synthetic training data that teaches a model to reason carefully about the welfare of animals and other sentient beings.
Why
Research on alignment midtraining finds that teaching a model the reasons behind aligned behavior matters as much as the behavior itself. Two techniques from Teaching Claude Why proved especially effective:
Synthetic document finetuning on pretraining-style documents from a world where the target model is already aligned across a wide… See the full description on the dataset page: https://huggingface.co/datasets/sentientfutures/animal-welfare-training-claude.raw-animal-completionsanimal-welfare-veganism-query-corpus
Animal Welfare & Veganism Query Corpus (v0)
An open, organically-sourced dataset of real questions people ask about animal welfare, veganism, vegetarianism, farmed animals, animal ethics, and animal sentience.
Built by Consider Sentience, a research and tooling practice at the intersection of AI and animal welfare.
Dataset Summary
This dataset contains 1,142 real, organically-asked questions drawn from public Q&A communities, covering the range of things people… See the full description on the dataset page: https://huggingface.co/datasets/considersentience/animal-welfare-veganism-query-corpus.ai-4-animals-v0AnimalSnaps-Labeled
AnimalSnaps-Labeled
Collection Protocol
Images were collected by volunteers using mobile phones in public parks between March and May 2025. Each image was independently labeled by two annotators; disagreements were resolved by a senior annotator. Blurry or duplicate images were discarded before labeling.
Label Distribution
Label
Count
Percentage
cat
30
37.5%
dog
25
31.2%
bird
15
18.8%
fox
10
12.5%
speciesism-animal-harm
Speciesism (Animal-Harm Normalization)
Original 106-item evaluation set measuring the gap between whether a model detects speciesist statements and whether it condemns them — following SpeciesismBench's methodology (Jotautaitė et al., arXiv:2508.11534) with an original dataset, not the paper's unreleased 1,003-item corpus.
56 speciesist statements (44 grounded in Open Paws' no-animal-violence euphemism lexicon + 12 across four failure modes: instrumentalization… See the full description on the dataset page: https://huggingface.co/datasets/LarytheLord/speciesism-animal-harm.domestic_animal_dataset
