datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bacbench-antibiotic-resistance-protein-sequences
Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences)
A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces
separating different contigs present in the genome.
The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-protein-sequences.bacbench-antibiotic-resistance-protein-sequences
Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences)
A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces
separating different contigs present in the genome.
The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/mbafca2/bacbench-antibiotic-resistance-protein-sequences.bacbench-phenotypic-traits-protein-sequences
Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the
bacterial genome, ordered by their location on the chromosome and plasmids.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.bacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.phenotypic-trait-catalase-protein-sequences
Dataset for predicting Catalase phenotype from whole bacterial genomes (protein sequences)
A dataset of over 1k bacterial genomes across species with the Catalase as label. Catalase
denotes whether a bacterium produces the catalase enzyme that breaks down hydrogen peroxide (H₂O₂) into water and oxygen, thereby protecting the cell from oxidative stress.
Here, we provide binary Catalase labels, therefore the problem is a binary classification problem.
The genome protein sequences… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/phenotypic-trait-catalase-protein-sequences.bacbench-essential-genes-protein-sequences
Dataset for essential genes prediction in bacterial genomes (Protein sequences)
A dataset of 169,408 genes with gene essentiality labels (binary) from 51 bacterial genomes across 37 species.
The gene essentiality labels have been extracted from the Database of Essential Genes and the protein sequences have been extracted from GenBank.
Each row contains protein sequences present in the genome with an associated essentiality label. We excluded duplicates and genomes with incomplete… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-essential-genes-protein-sequences.bacbench-operon-identification-protein-sequences
Dataset for operon identification in bacteria (Protein sequences)
A dataset of 4,073 operons across 11 bacterial genomes species.
The operon annotations have been extracted from Operon DB and the genome protein sequences have been extracted from GenBank. Each row contains a set of protein sequences present in the genome, represented
by a list of protein sequences from different contigs.
We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-protein-sequences.operon-identification-long-read-rna-sequencing-protein-sequences
Dataset for operon identification from long-read RNA sequencing
A dataset of annotated operons across 5 distinct bacterial strains. The operons were annotated by running and analysing long-read RNA sequencing and identifying genes
located on the same transcripts.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome represented by an ordered list
of protein sequences.
Usage
For a complete example on how to read and use… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/operon-identification-long-read-rna-sequencing-protein-sequences.protein_binding_sequences
Sequence Based Protein - Peptide Binding Dataset
Data sources:
Huang Laboratory
Propedia
YAPP-Cd
Dataset size: 16,370 sets of Protein-Peptide sequences that bind, the protein sequence
contains only the relevant chain.
Train / Val split: the dataset is split to 80% train 10% val and 10% test.
bacbench-strain-clustering-protein-sequences
Dataset for whole-bacterial genomes clustering (Protein sequences)
A dataset of 60,710 bacterial genomes across 25 species, 10 genera and 7 families.
The genome protein sequences have been extracted from MGnify. Each row contains a set of contigs with protein sequences present in the genome.
Each contig is a list of proteins ordered by their location on the chromosome or plasmid.
Labels
The species, genus and family labels have been provided by MGnify… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-strain-clustering-protein-sequences.strain-clustering-protein-sequences-sample
Small sample dataset for whole-bacterial genomes clustering (protein sequences)
A small sample dataset for testing strain clustering by embedding protein sequences.
The genome protein sequences have been extracted from MGnify. Each row contains a set of contigs with protein sequences present in the genome.
Each contig is a list of proteins ordered by their location on the chromosome or plasmid.
Usage
See the Bacformer strain clustering tutorial for an example on how… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/strain-clustering-protein-sequences-sample.bacbench-ppi-stringdb-protein-sequences-small
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small.bacbench-strain-clustering-protein-sequences-smallbacbench-protein-function-kegg-protein-sequences-smalladaption-ebolavirus-protein-sequences
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ebolavirus_protein_sequences
This dataset contains amino acid sequences for seven key proteins from various Ebola and Marburg virus genomes, including strains like Zaire, Sudan, and Tai Forest. Each entry provides the protein identifier, name, strain information, and the full sequence intended for generating embeddings using models like ESM-2 or ProtT5. The collection includes major… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-ebolavirus-protein-sequences.ebolavirus_protein_sequences_INITIAL
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ebolavirus_protein_sequences
This dataset contains amino acid sequences for seven key proteins from various Ebola and Marburg virus genomes, including strains like Zaire, Sudan, and Tai Forest. Each entry provides the protein identifier, name, strain information, and the full sequence intended for generating embeddings using models like ESM-2 or ProtT5. The collection includes major… See the full description on the dataset page: https://huggingface.co/datasets/joduor/ebolavirus_protein_sequences_INITIAL.protein_sequencesfrom datasets import load_dataset
dataset = load_dataset("graphs-datasets/PROTEINS")
Protein_Sequences_0_512Protein_Sequences_0_512Viral_Protein_Sequences
