Complexes
pdb_protein_ligand_complexes
How to use the data sets
This dataset contains about 36,000 unique pairs of protein sequences and ligand SMILES, and the coordinates
of their complexes from the PDB.
SMILES are assumed to be tokenized by the regex from P. Schwaller.
Ligand selection criteria
Only ligands
that have at least 3 atoms,
a molecular weight >= 100 Da,
and which are not among the 280 most common ligands in the PDB (this includes common additives like PEG, ADP, ..)
are considered.
Use… See the full description on the dataset page: https://huggingface.co/datasets/jglaser/pdb_protein_ligand_complexes.pdbbind_complexesA dataset to fine-tune language models on protein-ligand binding affinity and contact prediction.alphafold-complexes-metadataProtein complexes - May 2026
source files: https://ftp.ebi.ac.uk/pub/databases/alphafold/collaborations/nvda/
afdb-complexes-metadata-19milBoron-Complexes_SCXRD-dataset
Boron Complexes – Single-Crystal X-ray Diffraction Dataset
Dataset Summary
Boron-Complexes_SCXRD-dataset contains high-quality crystallographic data derived from single-crystal X-ray diffraction (SCXRD) analyses of a series of boron-based complexes. The dataset is organized into multiple files, each corresponding to an individual compound and including its refined crystallographic information in standard CIF format, along with associated structural and experimental… See the full description on the dataset page: https://huggingface.co/datasets/Ander18a/Boron-Complexes_SCXRD-dataset.pha_clustered_protein_complexesThis is a dataset of 10,000 interacting pairs of proteins obtained from UniProt,
and clustered using methods explain in this blog post. Note,
cluster 0 is over-represented in this dataset and this should be considered when creating train/test splits with this
data.
