datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pdb_protein_ligand_complexes
How to use the data sets
This dataset contains about 36,000 unique pairs of protein sequences and ligand SMILES, and the coordinates
of their complexes from the PDB.
SMILES are assumed to be tokenized by the regex from P. Schwaller.
Ligand selection criteria
Only ligands
that have at least 3 atoms,
a molecular weight >= 100 Da,
and which are not among the 280 most common ligands in the PDB (this includes common additives like PEG, ADP, ..)
are considered.
Use… See the full description on the dataset page: https://huggingface.co/datasets/jglaser/pdb_protein_ligand_complexes.pdbbind_complexesA dataset to fine-tune language models on protein-ligand binding affinity and contact prediction.alphafold-complexes-metadataProtein complexes - May 2026
source files: https://ftp.ebi.ac.uk/pub/databases/alphafold/collaborations/nvda/
afdb-complexes-metadata-19milBoron-Complexes_SCXRD-dataset
Boron Complexes – Single-Crystal X-ray Diffraction Dataset
Dataset Summary
Boron-Complexes_SCXRD-dataset contains high-quality crystallographic data derived from single-crystal X-ray diffraction (SCXRD) analyses of a series of boron-based complexes. The dataset is organized into multiple files, each corresponding to an individual compound and including its refined crystallographic information in standard CIF format, along with associated structural and experimental… See the full description on the dataset page: https://huggingface.co/datasets/Ander18a/Boron-Complexes_SCXRD-dataset.pha_clustered_protein_complexesThis is a dataset of 10,000 interacting pairs of proteins obtained from UniProt,
and clustered using methods explain in this blog post. Note,
cluster 0 is over-represented in this dataset and this should be considered when creating train/test splits with this
data.
Cationic_phenoxyimine_complexes_of_yttrium
Cite this dataset Oswald, A. D., Verrieux, L., Breuil, P. R., Olivier-Bourbigou, H., Thuilliez, J., Vaultier, F., Taoufik, M., Perrin, L., and Boisson, C. Cationic phenoxyimine complexes of yttrium. ColabFit, 2023. https://doi.org/10.60732/bd18acbe
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_y6bp7td4dle0_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Cationic_phenoxyimine_complexes_of_yttrium.ComplexesInformation about the dataset is detailed in the documentation:https://ai-chem.github.io/ChemX/overview/datasets_description.htmlYou can find the Croissant file in our GitHub repository:https://github.com/ai-chem/ChemX/tree/main/datasets/croissants
pha_clustered_protein_complexes_30K
Clustered Protein-Protein Complexes
This dataset is similar to AmelieSchreiber/pha_clustered_protein_complexes, but is larger. The methods used for clustering are the same,
with different hyperparameters. The threshold percentage for DBSCAN used to create the clusters is 0.35.
pha_clustered_protein_complexes_40K
