CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01heispv /protein_data_testsplit 1, 2 -> for sequences split 3, 4 -> for residues textn<1K0 likes2.2k downloads2y agoHugging Face02Metanova /Proteinstext1M<n<10M2 likes1.5k downloads1y agoHugging Face03heispv /protein_data_test_2textn<1K0 likes730 downloads2y agoHugging Face04genbio-ai /casp14-casp15-cameo-test-proteinstextn<1K0 likes630 downloads2y agoHugging Face05proteinea /secondary_structure_predictiontext10K<n<100K4 likes547 downloads4y agoHugging Face06MichelNivard /proteinLM-mixed-pretraining-v1 Pretraining mix for Protein language models In order to construct a solid pre-training data mixture for protein language models we sample a mix of proteins from 3 sources: MG_Prot50 (https://huggingface.co/datasets/tattabio/OMG_prot50): meta-genomic proteins created by clustering the Open MetaGenomic dataset (OMG) at 50% sequence identity. UniRef50: UniProt proteins from across all species clustered to 50% sequences identity, downloaded form UniProt on 31th of March 2025… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/proteinLM-mixed-pretraining-v1.texttoken-classification100M<n<1B0 likes354 downloads1y agoHugging Face07proteinea /solubilitytext10K<n<100K2 likes352 downloads4y agoHugging Face08genbio-ai /ProteinGYM-DMS-zeroshottabular1M<n<10M2 likes311 downloads2y agoHugging Face09proteinea /deeploctext1K<n<10K1 likes240 downloads4y agoHugging Face10PawanRamaMali /proteingym-fm-benchmark Protein Foundation Model Benchmark Results Zero-shot fitness prediction results for protein foundation models evaluated on the ProteinGym substitution benchmark (217 DMS assays, ~2.7M variants). Companion data for the paper: "From Sequence Encoders to Multimodal Systems: A Critical Survey of Protein Foundation Models" (IEEE TCBB 2026). Dataset configurations The dataset viewer exposes two configurations, because the files carry two different schemas that must not… See the full description on the dataset page: https://huggingface.co/datasets/PawanRamaMali/proteingym-fm-benchmark.tabular1M<n<10M0 likes239 downloads15d agoHugging Face11ChaoHou /protein_dynamic_properties Protein Sequences and Dynamic Properties for Training SeqDance and ESMDance This dataset contains protein sequences, dynamic properties, and feature weights used for training SeqDance and ESMDance, two protein language models designed to learn protein dynamic properties. training_test_data_sequence.csv This file contains 64,403 protein sequences with associated metadata for training and testing (excluding dynamicPDB). Columns: name – Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/ChaoHou/protein_dynamic_properties.text100K<n<1M2 likes230 downloads8mo agoHugging Face12proteinea /fluorescencetabular10K<n<100K1 likes226 downloads4y agoHugging Face13proteinea /remote_homologytabular10K<n<100K4 likes149 downloads4y agoHugging Face14Trelis /protein_stability_single_mutation Protein Data Stability - Single Mutation This repository contains data on the change in protein stability with a single mutation. Attribution of Data Sources Primary Source: Tsuboyama, K., Dauparas, J., Chen, J. et al. Mega-scale experimental analysis of protein folding stability in biology and design. Nature 620, 434–444 (2023). Link to the paper Dataset Link: Zenodo Record As to where the dataset comes from in this broader work, the relevant dataset (#3) is shown in… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/protein_stability_single_mutation.tabularquestion-answering100K<n<1M2 likes89 downloads3y agoHugging Face15monsoon-nlp /protein-pairs-uniprot-swissprot Protein Pairs and Similarity Selected protein similarities within training, test, and validation sets. Each protein gets two similarities selected at random and (usually) proteins within the top and bottom quintiles for similarity. The protein is represented by its UniProt ID and its amino acid sequence (using IUPAC-IUB codes where each amino acid maps to a letter of the alphabet, see: https://en.wikipedia.org/wiki/FASTA_format ). The distance column is cosine distance (identical =… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/protein-pairs-uniprot-swissprot.textsentence-similarity1M<n<10M2 likes66 downloads3y agoHugging Face16SaProtHub /Dataset-Structure_Class-ProteinShake Description Structural Class Prediction is a multi-class classification task to predict the correct structural class of given protein. This task is is built on the SCOP database. Protein Format: SA sequence (PDB) Splits The dataset is from ProteinShake Building datasets and benchmarks for deep learning on protein structures. We use the splits based on 70% structure similarity, with the number of training, validation and test set shown below: Train: 7990 Valid: 955 Test:… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Structure_Class-ProteinShake.text1K<n<10K0 likes64 downloads2y agoHugging Face17gcos /pyaptamer-proteins-shin2023text100K<n<1M0 likes60 downloads1y agoHugging Face18IAmKarthik /protein-compound-affinity-esm2-molformertext100K<n<1M0 likes59 downloads3mo agoHugging Face19AI4Protein /proteingym_87tabular1M<n<10M0 likes56 downloads10mo agoHugging Face20ronig /protein_binding_sequences Sequence Based Protein - Peptide Binding Dataset Data sources: Huang Laboratory Propedia YAPP-Cd Dataset size: 16,370 sets of Protein-Peptide sequences that bind, the protein sequence contains only the relevant chain. Train / Val split: the dataset is split to 80% train 10% val and 10% test. text10K<n<100K9 likes55 downloads3y agoHugging Face21monsoon-nlp /greenbeing-proteins GreenBeing Proteins dataset Proteins from UniProtKB (knowledge base), from select food crops and related species. Amino acid sequences use IUPAC-IUB codes where letters A-Z map to amino acids. Usage (due to different schema on splits): load_dataset("monsoon-nlp/greenbeing-proteins", "pretraining", split="pretraining") XML source from https://www.uniprot.org/help/downloads CoLab notebook: https://colab.research.google.com/drive/1M6sO0Ws6i5z9VUXIXopiOqo1OkQ7K-1g?usp=sharing… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/greenbeing-proteins.text1M<n<10M3 likes52 downloads2y agoHugging Face22lchen915 /proteinbase-af2af3-k5 proteinbase-af2af3-k5: AF2 + AF3 metrics with a K=5 AF3 diffusion ensemble Per-design AlphaFold2 and AlphaFold3 confidence metrics for the binder-target pairs in yk0/proteinbase_interactions, extended with a K=5 AlphaFold3 diffusion ensemble for uncertainty quantification. 2346 rows (one per binder-target pair); af3_k5_persample.csv holds all 5 samples per design. Two views (config selector in the viewer) ensemble (default): one row per design, with AF2, AF3… See the full description on the dataset page: https://huggingface.co/datasets/lchen915/proteinbase-af2af3-k5.tabular10K<n<100K0 likes52 downloads4mo agoHugging Face23AmelieSchreiber /family_split_protein_binding_sites UniProt Segmented Binding/Active Sites This is a train/test split of 209,571 protein sequences from UniProt of protein sequences with active sites and binding sites labels. All protein sequences (and corresponding binding site labels) are segmented into chunks of 1000 or less. Segmented sequences are indicated by a _partN suffix in the Entry column labels. The split is approximately 85/15 before segmentation. The proteins are sorted by family in decreasing order, with families… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/family_split_protein_binding_sites.texttoken-classification10K<n<100K0 likes42 downloads3y agoHugging Face24RosettaCommons /2J-Protein-Couplings2J-Protein-Coupling Dataset This data set was curated from the paper below accessed through the Biological Magnetic Resonance Data Bank (BMRB). There are a total of 3999 2J coupling taken from 5 different proteins and up to 10 different experiments. This dataset contains information regarding PDB ID, Sequence, 2J coupling data of 15N, 13C, and 1H. Data was curated and organized into this set of the five papers below, with the addition of the sequence taken from the Protein Data Bank. Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/2J-Protein-Couplings.tabularn<1K0 likes40 downloads6mo agoHugging Face25PDBEurope /protein_chain_conformational_states Schema description: The manually curated dataset of open-closed monomers is included here as benchmarking_monomeric_open_closed_conformers.csv. Column descriptions: Schema description: The manually curated dataset of open-closed monomers is included here as benchmarking_monomeric_open_closed_conformers.csv. Column descriptions: UNP_ACC | UniProt accession code UNP_START | Start of UniProt sequence for given PDBe entries UNP_END | End of UniProt sequence for given… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_chain_conformational_states.tabularfeature-extractionn<1K0 likes39 downloads3y agoHugging Face26SaProtHub /Dataset-Meta-scale-protein-stability Description This dataset contains ΔΔG value of various protein mutations (including some human designed proteins). The ΔΔG value indicates the stability of protein. Protein Format: AA sequence Splits traing: 486352 valid: 60995 test: 60492 Related paper Tsuboyama, K., Dauparas, J., Chen, J. et al. Mega-scale experimental analysis of protein folding stability in biology and design. Nature 620, 434–444 (2023). https://doi.org/10.1038/s41586-023-06328-6.… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Meta-scale-protein-stability.text100K<n<1M1 likes37 downloads2y agoHugging Face27AmelieSchreiber /aging_proteins Description of the Dataset This is (part of) the dataset used in Prediction and characterization of human ageing-related proteins by using machine learning. This can be used to train a binary sequence classifier using protein language models such as ESM-2. Please also see the github for the paper for more information. texttext-classification10K<n<100K2 likes36 downloads3y agoHugging Face28SaProtHub /Dataset-Structural_Similarity-ProteinShake Description Structure Similarity Prediction predicts the (aligned) Local Distance Difference Test (LDDT) of the structures given an unaligned pair of proteins. Target values are computed after alignment with TM-align for all pairs of 1000 randomly sampled single-chain proteins. Protein Format: SA sequence(PDB) Splits The dataset is from ProteinShake Building datasets and benchmarks for deep learning on protein structures. We use the splits based on 70% structure… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Structural_Similarity-ProteinShake.text100K<n<1M2 likes36 downloads2y agoHugging Face29yk0 /proteinbase_interactions ProteinBase Interactions Overview Each row represents a single binder-target pair with: the ProteinBase binder identifier the binder sequence the target name the target sequence an experimental binding-strength label a binary classification label design-class metadata an antibody flag indicating whether the binder is an antibody-derived design (Nanobody or scFv) Schema The dataset contains the following columns: proteinbase_id: Original ProteinBase… See the full description on the dataset page: https://huggingface.co/datasets/yk0/proteinbase_interactions.tabular1K<n<10K0 likes36 downloads4mo agoHugging Face30ClarusC64 /protein-mutation-stability-instability-v0.1 protein-mutation-stability-instability-v0.1 What this dataset does This dataset evaluates whether models can detect protein instability caused by mutation effects. Each row represents a simplified mutation scenario described through structural and interaction proxies. The task is to determine whether the mutation is likely to destabilize the protein. Core stability idea Mutation instability does not depend on mutation severity alone. A mutation may be tolerated… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/protein-mutation-stability-instability-v0.1.tabulartabular-classificationn<1K0 likes35 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.