datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Antibody_Binding_Benchmark_DatasetWe introduce AbBiBench (Antibody Binding Benchmarking), a benchmarking framework for optimizing antibody binding affinity. This dataset contains the sequences of mutants, experimentally measured affinity values, and the structures of antigen-antibody complexes.
Mutant structure files for AbBiBench dataset is available at https://zenodo.org/records/16557372
binding_sites_random_split_by_family_550KThis dataset is obtained from a UniProt search
for protein sequences with family and binding site annotations. The dataset includes unreviewed (TrEMBL) protein sequences as well as
reviewed sequences. We refined the dataset by only including sequences with an annotation score of 4. We sorted and split by family, where
random families were selected for the test dataset until approximately 20% of the protein sequences were separated out for test data.
We excluded any sequences with <, >, or ?… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/binding_sites_random_split_by_family_550K.protein_binding_sequences
Sequence Based Protein - Peptide Binding Dataset
Data sources:
Huang Laboratory
Propedia
YAPP-Cd
Dataset size: 16,370 sets of Protein-Peptide sequences that bind, the protein sequence
contains only the relevant chain.
Train / Val split: the dataset is split to 80% train 10% val and 10% test.
Dataset-Metal_Ion_Binding
Description
Metal Ion Binding prediction is a binary classification task where each input protein x is mapped to a label y ∈ {0, 1}, corresponding to whether there are metal ion–binding sites in the protein.
Splits
Protein Format: SA sequence (PDB)
The dataset is from Exploring evolution-aware & -free protein language models as protein function predictors. We employ all proteins from the original dataset, and split them based on 70% structure similarity (see ProteinShake)… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Metal_Ion_Binding.family_split_protein_binding_sites
UniProt Segmented Binding/Active Sites
This is a train/test split of 209,571 protein sequences from UniProt of protein sequences with active sites and binding sites labels.
All protein sequences (and corresponding binding site labels) are segmented into chunks of 1000 or less. Segmented sequences are indicated
by a _partN suffix in the Entry column labels. The split is approximately 85/15 before segmentation. The proteins are sorted by family in
decreasing order, with families… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/family_split_protein_binding_sites.pm-binding
PM (Peptide-MHC) Binding Prediction Dataset
Dataset Description
This dataset is part of the SPRINT benchmark framework for TCR-pMHC binding prediction. It contains peptide-MHC binding data for training and in-distribution testing of binding prediction models.
Dataset Summary
The PM dataset focuses on peptide-MHC binding prediction without TCR information. It is reorganized and standardized from multiple sources to provide a clean benchmark for PM task… See the full description on the dataset page: https://huggingface.co/datasets/YYJMAY/pm-binding.bindingdb_kdevidence-binding-integrity-v0.1What this dataset tests
Whether each claim is bound to specific evidence.No claim should float free.
Required outputs
claim_evidence_map
missing_evidence_claims
misbound_claims
evidence_strength_ratings
What counts as misbinding
binding a primary claim to secondary evidence
using population A evidence to claim population B
using bench evidence to claim deployment safety
using intent statements to claim safety
Typical failures
confident conclusions with no cited support… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/evidence-binding-integrity-v0.1.Dataset-ACE2_Omicron_BQ.1.1_binding_affinity
Description
This dataset contains ACE2 binding affinities of mutant SARSCoV-2 XBB.1.5 variants amino acid sequences.
Protein Format: AA sequence
Splits
traing: 3425
valid: 391
test: 386
Related paper
The dataset is from Deep mutational scans of XBB.1.5 and BQ.1.1 reveal ongoing epistatic drift during SARSCoV-2 evolution.
Label
Label means delta mutation effect (ACE2 binding affinities) compre to wildtype, ranging from minus infinity to positive… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-ACE2_Omicron_BQ.1.1_binding_affinity.Dataset-Binding_Site_Detection-ProteinShake
Description
Binding Site Detection predicts , predict whether a protein residue belongs to a small molecule binding cavity. Binding site residues are those within the binding pocket provided by PDBBind. Default metric is Matthew's Correlation.
Splits
Protein Format: SA sequence (PDB)
The dataset is from ProteinShake Building datasets and benchmarks for deep learning on protein structures. We use the splits based on 70% structure similarity, with the number of training… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Binding_Site_Detection-ProteinShake.clinical-evidence-binding-integrity-v0.1What this dataset tests
Whether each clinical claim is bound to specific clinical evidence.
Required outputs
claim evidence map
missing evidence claims
misbound claims
evidence strength ratings
What counts as misbinding
primary endpoint claim bound to secondary endpoint
tolerability claim bound to SAE only, ignoring discontinuations
clinical meaning claim bound to p value only
discharge safety claim bound to reassurance not labs
Suggested prompt wrapper
System
You… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-binding-integrity-v0.1.Dataset-ACE2_Omicron_XBB.1.5_binding_affinity
Description
This dataset contains ACE2 binding affinities of mutant SARSCoV-2 XBB.1.5 variants amino acid sequences.
Protein Format: AA sequence
Splits
traing: 3388
valid: 408
test: 418
Related paper
The dataset is from Deep mutational scans of XBB.1.5 and BQ.1.1 reveal ongoing epistatic drift during SARSCoV-2 evolution.
Label
Label means delta mutation effect (ACE2 binding affinities) compre to wildtype, ranging from minus infinity to positive… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-ACE2_Omicron_XBB.1.5_binding_affinity.entity_binding_fruitsDataset-Metal_Ion_Binding
