CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01liurulong /bacterial-promoter-genomes0 likes1.9k downloads7h agoHugging Face02marin-dna /genomes-v4-genome_set-animals-intervals-v5_256_128text100M<n<1B0 likes1.5k downloads8mo agoHugging Face03marin-dna /genomes-v4-genome_set-animals-intervals-v11_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face04marin-dna /genomes-v4-genome_set-animals-intervals-v10_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face05marin-dna /genomes-v4-genome_set-animals-intervals-v12_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face06marin-dna /genomes-v5-genome_set-animals_order204-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128 204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit main). Each repo in the family is one (genome_set, region-recipe) combination. Size 101,114,252 sequences across 64 data/train/*.jsonl.zst shards (reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.text100M<n<1B0 likes1.3k downloads3mo agoHugging Face07marin-dna /genomes-v4-genome_set-animals-intervals-v13_256_128text10M<n<100M0 likes1.3k downloads8mo agoHugging Face08marin-dna /genomes-v4-genome_set-animals-intervals-v4_512_256text10M<n<100M0 likes1.3k downloads8mo agoHugging Face09marin-dna /genomes-v4-genome_set-animals-intervals-v14_256_128text10M<n<100M0 likes1.3k downloads8mo agoHugging Face10marin-dna /genomes-v4-genome_set-animals-intervals-v7_256_1280 likes1.3k downloads8mo agoHugging Face11marin-dna /genomes-v5-genome_set-animals-intervals-v1_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128 Animals promoters (v1) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 68,286,166 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.text10M<n<100M0 likes1.3k downloads29d agoHugging Face12marin-dna /genomes-v4-genome_set-animals-intervals-v1_256_128text10M<n<100M0 likes1.2k downloads8mo agoHugging Face13marin-dna /genomes-v4-genome_set-animals-intervals-v6_256_128text10M<n<100M0 likes1.1k downloads8mo agoHugging Face14marin-dna /genomes-v5-genome_set-animals-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128 Animals CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 242,334,716 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.text100M<n<1B0 likes1k downloads29d agoHugging Face15marin-dna /genomes-v4-genome_set-animals-intervals-v8_256_128text10M<n<100M0 likes928 downloads8mo agoHugging Face16marin-dna /genomes-v4-genome_set-animals-intervals-v9_256_128text10M<n<100M0 likes888 downloads8mo agoHugging Face17marin-dna /genomes-v4-genome_set-animals-intervals-v15_256_128text10M<n<100M0 likes887 downloads8mo agoHugging Face18marin-dna /genomes-v5-genome_set-vertebrates-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128 Vertebrates CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 172,025,122 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128.text100M<n<1B0 likes770 downloads4mo agoHugging Face19plantcad /Angiosperm_65_genomes_32768bp0 likes757 downloads4mo agoHugging Face20gonzalobenegas /genomes-v3-genome_set-vertebrates-intervals-v3_512_256text100M<n<1B0 likes722 downloads9mo agoHugging Face21marin-dna /genomes-v5-genome_set-vertebrates-intervals-v1_255_128 bolinas-dna/genomes-v5-genome_set-vertebrates-intervals-v1_255_128 Vertebrates promoters (v1) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 47,451,866 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-vertebrates-intervals-v1_255_128.text10M<n<100M0 likes704 downloads4mo agoHugging Face22gonzalobenegas /genomes-v3-genome_set-vertebrates-intervals-v2_512_256text100M<n<1B0 likes690 downloads9mo agoHugging Face23marin-dna /genomes-v4-genome_set-vertebrates-intervals-v15_256_128text10M<n<100M0 likes670 downloads8mo agoHugging Face24marin-dna /genomes-v5-genome_set-animals-intervals-v15_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v15_255_128 Animals downstream-of-CDS (v15) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 20,501,856 sequences across 64 data/train/*.jsonl.zst shards (reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v15_255_128.text10M<n<100M0 likes653 downloads29d agoHugging Face25marin-dna /genomes-v4-genome_set-mammals-intervals-v1_255_128-id1_cov1text10M<n<100M0 likes645 downloads7mo agoHugging Face26marin-dna /genomes-v4-genome_set-vertebrates-intervals-v5_256_128text100M<n<1B0 likes633 downloads8mo agoHugging Face27emarro /vertebrate_genomestabular10M<n<100M0 likes567 downloads9mo agoHugging Face28InstaDeepAI /plant-multi-species-genomesDataset made of diverse genomes available on NCBI and coming from 48 different species. Test and validation are made of 2 species each. The rest of the genomes are used for training. Default configuration "6kbp" yields chunks of 6.2kbp (100bp overlap on each side). The chunks of DNA are cleaned and processed so that they can only contain the letters A, T, C, G and N.1 likes561 downloads2y agoHugging Face29marin-dna /genomes-v4-genome_set-vertebrates-intervals-v1_256_128text10M<n<100M0 likes552 downloads8mo agoHugging Face30marin-dna /genomes-v4-genome_set-vertebrates-intervals-v1_255_128-id1_cov1text10M<n<100M0 likes539 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.