CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01macwiatrak /bacbench-antibiotic-resistance-protein-sequences Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences) A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces separating different contigs present in the genome. The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-protein-sequences.text10K<n<100K0 likes881 downloads1y agoHugging Face02mbafca2 /bacbench-antibiotic-resistance-protein-sequences Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences) A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces separating different contigs present in the genome. The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/mbafca2/bacbench-antibiotic-resistance-protein-sequences.text10K<n<100K0 likes436 downloads7mo agoHugging Face03macwiatrak /bacbench-phenotypic-traits-protein-sequences Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences) A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels. The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the bacterial genome, ordered by their location on the chromosome and plasmids. The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.text10K<n<100K0 likes297 downloads10mo agoHugging Face04macwiatrak /bacbench-ppi-stringdb-protein-sequences Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.tabular1K<n<10K0 likes262 downloads1y agoHugging Face05macwiatrak /phenotypic-trait-catalase-protein-sequences Dataset for predicting Catalase phenotype from whole bacterial genomes (protein sequences) A dataset of over 1k bacterial genomes across species with the Catalase as label. Catalase denotes whether a bacterium produces the catalase enzyme that breaks down hydrogen peroxide (H₂O₂) into water and oxygen, thereby protecting the cell from oxidative stress. Here, we provide binary Catalase labels, therefore the problem is a binary classification problem. The genome protein sequences… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/phenotypic-trait-catalase-protein-sequences.text1K<n<10K0 likes94 downloads1y agoHugging Face06macwiatrak /bacbench-essential-genes-protein-sequences Dataset for essential genes prediction in bacterial genomes (Protein sequences) A dataset of 169,408 genes with gene essentiality labels (binary) from 51 bacterial genomes across 37 species. The gene essentiality labels have been extracted from the Database of Essential Genes and the protein sequences have been extracted from GenBank. Each row contains protein sequences present in the genome with an associated essentiality label. We excluded duplicates and genomes with incomplete… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-essential-genes-protein-sequences.textn<1K0 likes86 downloads5mo agoHugging Face07macwiatrak /bacbench-operon-identification-protein-sequences Dataset for operon identification in bacteria (Protein sequences) A dataset of 4,073 operons across 11 bacterial genomes species. The operon annotations have been extracted from Operon DB and the genome protein sequences have been extracted from GenBank. Each row contains a set of protein sequences present in the genome, represented by a list of protein sequences from different contigs. We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-protein-sequences.textn<1K0 likes85 downloads1y agoHugging Face08macwiatrak /operon-identification-long-read-rna-sequencing-protein-sequences Dataset for operon identification from long-read RNA sequencing A dataset of annotated operons across 5 distinct bacterial strains. The operons were annotated by running and analysing long-read RNA sequencing and identifying genes located on the same transcripts. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome represented by an ordered list of protein sequences. Usage For a complete example on how to read and use… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/operon-identification-long-read-rna-sequencing-protein-sequences.textn<1K0 likes63 downloads1y agoHugging Face09ronig /protein_binding_sequences Sequence Based Protein - Peptide Binding Dataset Data sources: Huang Laboratory Propedia YAPP-Cd Dataset size: 16,370 sets of Protein-Peptide sequences that bind, the protein sequence contains only the relevant chain. Train / Val split: the dataset is split to 80% train 10% val and 10% test. text10K<n<100K9 likes55 downloads3y agoHugging Face10macwiatrak /bacbench-strain-clustering-protein-sequences Dataset for whole-bacterial genomes clustering (Protein sequences) A dataset of 60,710 bacterial genomes across 25 species, 10 genera and 7 families. The genome protein sequences have been extracted from MGnify. Each row contains a set of contigs with protein sequences present in the genome. Each contig is a list of proteins ordered by their location on the chromosome or plasmid. Labels The species, genus and family labels have been provided by MGnify… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-strain-clustering-protein-sequences.text10K<n<100K0 likes49 downloads1y agoHugging Face11macwiatrak /strain-clustering-protein-sequences-sample Small sample dataset for whole-bacterial genomes clustering (protein sequences) A small sample dataset for testing strain clustering by embedding protein sequences. The genome protein sequences have been extracted from MGnify. Each row contains a set of contigs with protein sequences present in the genome. Each contig is a list of proteins ordered by their location on the chromosome or plasmid. Usage See the Bacformer strain clustering tutorial for an example on how… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/strain-clustering-protein-sequences-sample.textn<1K0 likes48 downloads1y agoHugging Face12macwiatrak /bacbench-ppi-stringdb-protein-sequences-small Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small.textn<1K0 likes47 downloads5mo agoHugging Face13macwiatrak /bacbench-strain-clustering-protein-sequences-smalltext1K<n<10K0 likes33 downloads10mo agoHugging Face14macwiatrak /bacbench-protein-function-kegg-protein-sequences-smalltextn<1K0 likes26 downloads10mo agoHugging Face15joduor /adaption-ebolavirus-protein-sequences This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-ebolavirus_protein_sequences This dataset contains amino acid sequences for seven key proteins from various Ebola and Marburg virus genomes, including strains like Zaire, Sudan, and Tai Forest. Each entry provides the protein identifier, name, strain information, and the full sequence intended for generating embeddings using models like ESM-2 or ProtT5. The collection includes major… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-ebolavirus-protein-sequences.tabular1K<n<10K0 likes7 downloads4mo agoHugging Face16joduor /ebolavirus_protein_sequences_INITIAL This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-ebolavirus_protein_sequences This dataset contains amino acid sequences for seven key proteins from various Ebola and Marburg virus genomes, including strains like Zaire, Sudan, and Tai Forest. Each entry provides the protein identifier, name, strain information, and the full sequence intended for generating embeddings using models like ESM-2 or ProtT5. The collection includes major… See the full description on the dataset page: https://huggingface.co/datasets/joduor/ebolavirus_protein_sequences_INITIAL.tabular1K<n<10K0 likes7 downloads4mo agoHugging Face17gsspdev /protein_sequencesfrom datasets import load_dataset dataset = load_dataset("graphs-datasets/PROTEINS") 0 likes6 downloads3y agoHugging Face18JuIm /Protein_Sequences_0_512text100K<n<1M0 likes6 downloads2y agoHugging Face19Mowriss /Protein_Sequences_0_512text100K<n<1M0 likes5 downloads5mo agoHugging Face20cdechristopher /Viral_Protein_Sequences0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.