CoolFace
Datasetpublic

macwiatrak/strain-clustering-protein-sequences-sample

Small sample dataset for whole-bacterial genomes clustering (protein sequences) A small sample dataset for testing strain clustering by embedding protein sequences. The genome protein sequences have been extracted from MGnify. Each row contains a set of contigs with protein sequences present in the genome. Each contig is a list of proteins ordered by their location on the chromosome or plasmid. Usage See the Bacformer strain clustering tutorial for an example on… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/strain-clustering-protein-sequences-sample.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes33downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
macwiatrak/strain-clustering-protein-sequences-sample · CoolFace