CoolFace
Datasetpublic

macwiatrak/strain-clustering-protein-sequences-sample

Small sample dataset for whole-bacterial genomes clustering (protein sequences) A small sample dataset for testing strain clustering by embedding protein sequences. The genome protein sequences have been extracted from MGnify. Each row contains a set of contigs with protein sequences present in the genome. Each contig is a list of proteins ordered by their location on the chromosome or plasmid. Usage See the Bacformer strain clustering tutorial for an example on… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/strain-clustering-protein-sequences-sample.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes33downloads
settings

This repository belongs to macwiatrak on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namestrain-clustering-protein-sequences-sample
visibilitypublic
licenceapache-2.0
gatedno
ownermacwiatrak
Account settings
macwiatrak/strain-clustering-protein-sequences-sample · CoolFace