CoolFace
Datasetpublic

macwiatrak/bacbench-strain-clustering-protein-sequences

Dataset for whole-bacterial genomes clustering (Protein sequences) A dataset of 60,710 bacterial genomes across 25 species, 10 genera and 7 families. The genome protein sequences have been extracted from MGnify. Each row contains a set of contigs with protein sequences present in the genome. Each contig is a list of proteins ordered by their location on the chromosome or plasmid. Labels The species, genus and family labels have been provided by MGnify… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-strain-clustering-protein-sequences.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes57downloads
5 commits on main
9cdb7011y ago

Update README.md

macwiatrak
fc5169f1y ago

Update README.md

macwiatrak
320eda71y ago

Upload dataset (part 00001-of-00002)

macwiatrak
30c8c651y ago

Upload dataset (part 00000-of-00002)

macwiatrak
fa5e4301y ago

initial commit

macwiatrak