datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genomes-v4-genome_set-mammals-intervals-v1_255_128-id1_cov1genomes-v4-genome_set-mammals-intervals-v16_254_127-id0.3_cov0.3genomes-v2-genome_set-mammals-intervals-v2_512_256genomes-v4-genome_set-mammals-intervals-v5_256_128genomes-v5-genome_set-mammals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-mammals-intervals-v5_255_128
Mammals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
41,848,032 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals-intervals-v5_255_128.genomes-v4-genome_set-mammals-intervals-v1_256_128vertebrate-v1-cds_mammals_only
marin-dna/vertebrate-v1-cds_mammals_only
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment. This
draft covers the cds region cohort with mammals_only species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source non-repeat-masked sequence, and… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-cds_mammals_only.genomes-v5-genome_set-mammals-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-mammals-intervals-v1_255_128
Mammals promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
12,926,544 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals-intervals-v1_255_128.genomes-v4-genome_set-mammals-intervals-v15_256_128genomes-v2-genome_set-mammals-intervals-v1_512_256genomes-v5-genome_set-mammals_seg20-intervals-v32_255_128
bolinas-dna/genomes-v5-genome_set-mammals_seg20-intervals-v32_255_128
20 mammals projected phastCons enhancers (v32) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
6,648,120 sequences across 64 data/train/*.jsonl.zst shards
(reverse… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals_seg20-intervals-v32_255_128.genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128
bolinas-dna/genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128
20 mammals (segmentation) segmentation enhancers (v20) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
8,672,102 sequences across 64… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128.genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128
bolinas-dna/genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128
20 mammals projected conserved enhancers (v30) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
6,549,730 sequences across 64 data/train/*.jsonl.zst shards
(reverse… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128.genomes-v5-genome_set-mammals_seg20-intervals-v33_255_128
bolinas-dna/genomes-v5-genome_set-mammals_seg20-intervals-v33_255_128
20 mammals projected phastCons enhancers (≥50 bp) (v33) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
4,791,766 sequences across 64 data/train/*.jsonl.zst shards… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals_seg20-intervals-v33_255_128.genomes-v5-genome_set-mammals_seg20-intervals-v31_255_128
bolinas-dna/genomes-v5-genome_set-mammals_seg20-intervals-v31_255_128
20 mammals projected conserved enhancers (≥50 bp) (v31) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
4,112,634 sequences across 64 data/train/*.jsonl.zst shards… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals_seg20-intervals-v31_255_128.mammals_32kgenomes-v4-genome_set-mammals-intervals-v1_255_128-id0.3_cov0.3genomes-v5-genome_set-mammals-intervals-v15_255_128
bolinas-dna/genomes-v5-genome_set-mammals-intervals-v15_255_128
Mammals downstream-of-CDS (v15) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
3,801,906 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals-intervals-v15_255_128.ref_seq_mammals_part_1mammals_16kgenomes-v3-genome_set-mammals-intervals-v3_512_256ref_seq_mammals_part_2watkins_marine_mammals_fullgenomes-v3-genome_set-mammals-intervals-v1_512_256genomes-v3-genome_set-mammals-intervals-v2_512_256Termcat_Marine_Mammals
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/lcr/19294
Description
Marine Mammals terms
Citation
Termcat Marine Mammals (2022). Version unspecified. [Dataset (Lexical/Conceptual Resource)]. Source: European Language Grid. https://live.european-language-grid.eu/catalogue/lcr/19294
mammals_32k_tokenizedgenomes-mammals-balanced-v1genomes-mammals-balanced-v1-1024test_annotated_mammals
