macwiatrak/strain-clustering-protein-sequences-sample
Small sample dataset for whole-bacterial genomes clustering (protein sequences) A small sample dataset for testing strain clustering by embedding protein sequences. The genome protein sequences have been extracted from MGnify. Each row contains a set of contigs with protein sequences present in the genome. Each contig is a list of proteins ordered by their location on the chromosome or plasmid. Usage See the Bacformer strain clustering tutorial for an example on… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/strain-clustering-protein-sequences-sample.
Small sample dataset for whole-bacterial genomes clustering (protein sequences)
A small sample dataset for testing strain clustering by embedding protein sequences.
The genome protein sequences have been extracted from MGnify. Each row contains a set of contigs with protein sequences present in the genome. Each contig is a list of proteins ordered by their location on the chromosome or plasmid.
Usage
See the Bacformer strain clustering tutorial for an example on how to use the dataset.
