datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
protein_data_testsplit 1, 2 -> for sequences
split 3, 4 -> for residues
Proteinsprotein_data_test_2casp14-casp15-cameo-test-proteinssecondary_structure_predictionproteinLM-mixed-pretraining-v1
Pretraining mix for Protein language models
In order to construct a solid pre-training data mixture for protein language models we sample a mix of proteins from 3 sources:
MG_Prot50 (https://huggingface.co/datasets/tattabio/OMG_prot50): meta-genomic proteins created by clustering the Open MetaGenomic dataset (OMG) at 50% sequence identity.
UniRef50: UniProt proteins from across all species clustered to 50% sequences identity, downloaded form UniProt on 31th of March 2025… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/proteinLM-mixed-pretraining-v1.solubilityProteinGYM-DMS-zeroshotdeeplocproteingym-fm-benchmark
Protein Foundation Model Benchmark Results
Zero-shot fitness prediction results for protein foundation models evaluated on
the ProteinGym substitution benchmark (217 DMS
assays, ~2.7M variants).
Companion data for the paper: "From Sequence Encoders to Multimodal Systems:
A Critical Survey of Protein Foundation Models" (IEEE TCBB 2026).
Dataset configurations
The dataset viewer exposes two configurations, because the files carry two
different schemas that must not… See the full description on the dataset page: https://huggingface.co/datasets/PawanRamaMali/proteingym-fm-benchmark.protein_dynamic_properties
Protein Sequences and Dynamic Properties for Training SeqDance and ESMDance
This dataset contains protein sequences, dynamic properties, and feature weights used for training SeqDance and ESMDance, two protein language models designed to learn protein dynamic properties.
training_test_data_sequence.csv
This file contains 64,403 protein sequences with associated metadata for training and testing (excluding dynamicPDB).
Columns:
name – Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/ChaoHou/protein_dynamic_properties.fluorescenceremote_homologyprotein_stability_single_mutation
Protein Data Stability - Single Mutation
This repository contains data on the change in protein stability with a single mutation.
Attribution of Data Sources
Primary Source: Tsuboyama, K., Dauparas, J., Chen, J. et al. Mega-scale experimental analysis of protein folding stability in biology and design. Nature 620, 434–444 (2023). Link to the paper
Dataset Link: Zenodo Record
As to where the dataset comes from in this broader work, the relevant dataset (#3) is shown in… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/protein_stability_single_mutation.protein-pairs-uniprot-swissprot
Protein Pairs and Similarity
Selected protein similarities within training, test, and validation sets.
Each protein gets two similarities selected at random and (usually) proteins within the top and bottom quintiles for similarity.
The protein is represented by its UniProt ID and its amino acid sequence (using IUPAC-IUB codes where each amino acid maps to a letter of the alphabet, see: https://en.wikipedia.org/wiki/FASTA_format ).
The distance column is cosine distance (identical =… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/protein-pairs-uniprot-swissprot.Dataset-Structure_Class-ProteinShake
Description
Structural Class Prediction is a multi-class classification task to predict the correct structural class of given protein. This task is is built on the SCOP database.
Protein Format: SA sequence (PDB)
Splits
The dataset is from ProteinShake Building datasets and benchmarks for deep learning on protein structures. We use the splits based on 70% structure similarity, with the number of training, validation and test set shown below:
Train: 7990
Valid: 955
Test:… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Structure_Class-ProteinShake.pyaptamer-proteins-shin2023protein-compound-affinity-esm2-molformerproteingym_87protein_binding_sequences
Sequence Based Protein - Peptide Binding Dataset
Data sources:
Huang Laboratory
Propedia
YAPP-Cd
Dataset size: 16,370 sets of Protein-Peptide sequences that bind, the protein sequence
contains only the relevant chain.
Train / Val split: the dataset is split to 80% train 10% val and 10% test.
greenbeing-proteins
GreenBeing Proteins dataset
Proteins from UniProtKB (knowledge base), from select food crops and related species.
Amino acid sequences use IUPAC-IUB codes where letters A-Z map to amino acids.
Usage (due to different schema on splits):
load_dataset("monsoon-nlp/greenbeing-proteins", "pretraining", split="pretraining")
XML source from https://www.uniprot.org/help/downloads
CoLab notebook: https://colab.research.google.com/drive/1M6sO0Ws6i5z9VUXIXopiOqo1OkQ7K-1g?usp=sharing… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/greenbeing-proteins.proteinbase-af2af3-k5
proteinbase-af2af3-k5: AF2 + AF3 metrics with a K=5 AF3 diffusion ensemble
Per-design AlphaFold2 and AlphaFold3 confidence metrics for the binder-target
pairs in yk0/proteinbase_interactions, extended with a
K=5 AlphaFold3 diffusion ensemble for uncertainty quantification.
2346 rows (one per binder-target pair); af3_k5_persample.csv holds all 5
samples per design.
Two views (config selector in the viewer)
ensemble (default): one row per design, with AF2, AF3… See the full description on the dataset page: https://huggingface.co/datasets/lchen915/proteinbase-af2af3-k5.family_split_protein_binding_sites
UniProt Segmented Binding/Active Sites
This is a train/test split of 209,571 protein sequences from UniProt of protein sequences with active sites and binding sites labels.
All protein sequences (and corresponding binding site labels) are segmented into chunks of 1000 or less. Segmented sequences are indicated
by a _partN suffix in the Entry column labels. The split is approximately 85/15 before segmentation. The proteins are sorted by family in
decreasing order, with families… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/family_split_protein_binding_sites.2J-Protein-Couplings2J-Protein-Coupling Dataset
This data set was curated from the paper below accessed through the Biological Magnetic Resonance Data Bank (BMRB). There are a total of 3999 2J coupling taken from 5 different proteins and up to 10 different experiments. This dataset contains information regarding PDB ID, Sequence, 2J coupling data of 15N, 13C, and 1H. Data was curated and organized into this set of the five papers below, with the addition of the sequence taken from the Protein Data Bank.
Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/2J-Protein-Couplings.protein_chain_conformational_states
Schema description:
The manually curated dataset of open-closed monomers is included here as benchmarking_monomeric_open_closed_conformers.csv.
Column descriptions:
Schema description:
The manually curated dataset of open-closed monomers is included here as benchmarking_monomeric_open_closed_conformers.csv.
Column descriptions:
UNP_ACC | UniProt accession code
UNP_START | Start of UniProt sequence for given PDBe entries
UNP_END | End of UniProt sequence for given… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_chain_conformational_states.Dataset-Meta-scale-protein-stability
Description
This dataset contains ΔΔG value of various protein mutations (including some human designed proteins). The ΔΔG value indicates the stability of protein.
Protein Format: AA sequence
Splits
traing: 486352
valid: 60995
test: 60492
Related paper
Tsuboyama, K., Dauparas, J., Chen, J. et al. Mega-scale experimental analysis of protein folding stability in biology and design. Nature 620, 434–444 (2023).
https://doi.org/10.1038/s41586-023-06328-6.… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Meta-scale-protein-stability.aging_proteins
Description of the Dataset
This is (part of) the dataset used in
Prediction and characterization of human ageing-related proteins by using machine learning.
This can be used to train a binary sequence classifier using protein language models such as ESM-2.
Please also see the github for the paper for more information.
Dataset-Structural_Similarity-ProteinShake
Description
Structure Similarity Prediction predicts the (aligned) Local Distance Difference Test (LDDT) of the structures given an unaligned pair of proteins. Target values are computed after alignment with TM-align for all pairs of 1000 randomly sampled single-chain proteins.
Protein Format: SA sequence(PDB)
Splits
The dataset is from ProteinShake Building datasets and benchmarks for deep learning on protein structures. We use the splits based on 70% structure… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Structural_Similarity-ProteinShake.proteinbase_interactions
ProteinBase Interactions
Overview
Each row represents a single binder-target pair with:
the ProteinBase binder identifier
the binder sequence
the target name
the target sequence
an experimental binding-strength label
a binary classification label
design-class metadata
an antibody flag indicating whether the binder is an antibody-derived design (Nanobody or scFv)
Schema
The dataset contains the following columns:
proteinbase_id: Original ProteinBase… See the full description on the dataset page: https://huggingface.co/datasets/yk0/proteinbase_interactions.protein-mutation-stability-instability-v0.1
protein-mutation-stability-instability-v0.1
What this dataset does
This dataset evaluates whether models can detect protein instability caused by mutation effects.
Each row represents a simplified mutation scenario described through structural and interaction proxies.
The task is to determine whether the mutation is likely to destabilize the protein.
Core stability idea
Mutation instability does not depend on mutation severity alone.
A mutation may be tolerated… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/protein-mutation-stability-instability-v0.1.
