Protein Sequences
bacbench-antibiotic-resistance-protein-sequences
Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences)
A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces
separating different contigs present in the genome.
The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-protein-sequences.bacbench-antibiotic-resistance-protein-sequences
Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences)
A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces
separating different contigs present in the genome.
The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/mbafca2/bacbench-antibiotic-resistance-protein-sequences.bacbench-phenotypic-traits-protein-sequences
Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the
bacterial genome, ordered by their location on the chromosome and plasmids.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.bacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.phenotypic-trait-catalase-protein-sequences
Dataset for predicting Catalase phenotype from whole bacterial genomes (protein sequences)
A dataset of over 1k bacterial genomes across species with the Catalase as label. Catalase
denotes whether a bacterium produces the catalase enzyme that breaks down hydrogen peroxide (H₂O₂) into water and oxygen, thereby protecting the cell from oxidative stress.
Here, we provide binary Catalase labels, therefore the problem is a binary classification problem.
The genome protein sequences… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/phenotypic-trait-catalase-protein-sequences.bacbench-essential-genes-protein-sequences
Dataset for essential genes prediction in bacterial genomes (Protein sequences)
A dataset of 169,408 genes with gene essentiality labels (binary) from 51 bacterial genomes across 37 species.
The gene essentiality labels have been extracted from the Database of Essential Genes and the protein sequences have been extracted from GenBank.
Each row contains protein sequences present in the genome with an associated essentiality label. We excluded duplicates and genomes with incomplete… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-essential-genes-protein-sequences.
