datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bacterial-promoter-genomesgenomes-v4-genome_set-animals-intervals-v5_256_128genomes-v4-genome_set-animals-intervals-v11_256_128genomes-v4-genome_set-animals-intervals-v10_256_128genomes-v4-genome_set-animals-intervals-v12_256_128genomes-v5-genome_set-animals_order204-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128
204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
main). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
101,114,252 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v13_256_128genomes-v4-genome_set-animals-intervals-v4_512_256genomes-v4-genome_set-animals-intervals-v14_256_128genomes-v4-genome_set-animals-intervals-v7_256_128genomes-v5-genome_set-animals-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128
Animals promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
68,286,166 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.genomes-v4-genome_set-animals-intervals-v1_256_128genomes-v4-genome_set-animals-intervals-v6_256_128genomes-v5-genome_set-animals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128
Animals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
242,334,716 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v8_256_128genomes-v4-genome_set-animals-intervals-v9_256_128genomes-v4-genome_set-animals-intervals-v15_256_128genomes-v5-genome_set-vertebrates-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128
Vertebrates CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
172,025,122 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128.Angiosperm_65_genomes_32768bpgenomes-v3-genome_set-vertebrates-intervals-v3_512_256genomes-v5-genome_set-vertebrates-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-vertebrates-intervals-v1_255_128
Vertebrates promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
47,451,866 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-vertebrates-intervals-v1_255_128.genomes-v3-genome_set-vertebrates-intervals-v2_512_256genomes-v4-genome_set-vertebrates-intervals-v15_256_128genomes-v5-genome_set-animals-intervals-v15_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v15_255_128
Animals downstream-of-CDS (v15) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
20,501,856 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v15_255_128.genomes-v4-genome_set-mammals-intervals-v1_255_128-id1_cov1genomes-v4-genome_set-vertebrates-intervals-v5_256_128vertebrate_genomesplant-multi-species-genomesDataset made of diverse genomes available on NCBI and coming from 48 different species.
Test and validation are made of 2 species each. The rest of the genomes are used for training.
Default configuration "6kbp" yields chunks of 6.2kbp (100bp overlap on each side). The chunks of DNA are cleaned and processed so that
they can only contain the letters A, T, C, G and N.genomes-v4-genome_set-vertebrates-intervals-v1_256_128genomes-v4-genome_set-vertebrates-intervals-v1_255_128-id1_cov1
