CoolFace
Datasetpublic

marin-dna/genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128

bolinas-dna/genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128 20 mammals projected conserved enhancers (v30) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 6,549,730 sequences across 64 data/train/*.jsonl.zst shards (reverse… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes180downloads
4 commits on main
0175a974mo ago

Add dataset card (README.md)

gonzalobenegas
20c02ba5mo ago

Add files using upload-large-folder tool

gonzalobenegas
b8f10fc5mo ago

Add files using upload-large-folder tool

gonzalobenegas
cd0c0f25mo ago

initial commit

gonzalobenegas