datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SAAINTDB
SAAINTDB
This dataset is a curated version of the SAAINT-DB converted into a format compatible with the Hugging Face Datasets for machine learning applications.
The dataset contains 21,400 antibody entries derived from 11,304 PDB structures, reflecting the available structures as of February 2026. Each entry corresponds to an antibody chain and is uniquely identified using the PDB_ID_chain field (PDB ID + chain ID).
Dataset Splits
The dataset was split at the PDB… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAAINTDB.SAbDab
ML Application Curated SAbDab
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the cache directory, e.g.
$ HF_HOME=${HOME}/.cache/huggingface/
$ export HF_HOME
then, from within python load the datasets library
>>> import datasets
Load model datasets
To load… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab.AfCycDesign
Dataset Card for AfCycDesign
Hallucinated scaffolds used by AfCycDesign for cyclic peptide design.
Dataset Details
Sets 7-16 of hallucinated peptide cif files and experimental CCDC structures.
Dataset Description
This dataset contains hallucinated cyclic peptide scaffold structures (in CIF format) generated using AfCycDesign, a deep learning approach built on AlphaFold2 for de novo design of cyclic peptides. The scaffolds span peptide lengths of 7–16 residues… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/AfCycDesign.FPbase
FPbase: The Fluorescent Protein Database
FPbase is a free, open-source, community-editable database of fluorescent proteins and their properties, aimed at aggregating structured, searchable information useful to the imaging community and FP developers. Visit fpbase.org for more.
This dataset updated on ,March 1st, 2026, collects FPbase fluorescent protein records (e.g., names, identifiers, sequences, and photophysical properties) for downstream analysis and modeling.… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/FPbase.NAKBOriginal Paper:
Lawson CL, Berman HM, Vallat B, Chen L, Zirbel C (2024) The Nucleic Acid Knowledgebase: a new portal for 3D structural information about nucleic acids. Nucleic Acids Research 52, D245-D254.
https://doi.org/10.1093/nar/gkad957
Nucleic Acid Knowledgebase (NAKB)
NAKB data set contains 21166 structures including Nucleic Acids, Protein, and Ligand Annotations, and determined 3D structures found in the Nucleic Acid Database (NDB) and the Protein Data Bank (PDB), including… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/NAKB.PTMint
PTMint
This dataset is derived from PTMint (https://ptmint.sjtu.edu.cn/), (Post Translational Modifications that are associated with Protein-Protein Interactions) that contains manually curated complete experimental evidence of the PTM effecting on protein-protein interactions in multiple organisms, including H. sapines, A. thaliana, C. elegans, D. melanogaster, S. cerevisiae and S. pombe.
This Hugging Face dataset repository provides PTMint-derived tables including a precomputed… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/PTMint.CatPred-DB
CatPred-DB: Enzyme Kinetic Parameters Database
Paper: CatPred: A comprehensive framework for deep learning in vitro enzyme kinetic parameters
GitHub: https://github.com/maranasgroup/CatPred-DB
Dataset Description
CatPred-DB contains the benchmark datasets introduced alongside the CatPred deep learning framework for predicting in vitro enzyme kinetic parameters. The datasets cover three key kinetic parameters:
Parameter
Description
Datapoints
kcat
Turnover… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/CatPred-DB.2J-Protein-Couplings2J-Protein-Coupling Dataset
This data set was curated from the paper below accessed through the Biological Magnetic Resonance Data Bank (BMRB). There are a total of 3999 2J coupling taken from 5 different proteins and up to 10 different experiments. This dataset contains information regarding PDB ID, Sequence, 2J coupling data of 15N, 13C, and 1H. Data was curated and organized into this set of the five papers below, with the addition of the sequence taken from the Protein Data Bank.
Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/2J-Protein-Couplings.BELKA-DEL-Experimental-BenchmarkData
BELKA-DEL-Experimental-BenchmarkData
This dataset comprises a curated collection of PDB structures, designed as an experimental validation benchmark for models trained on the Big Encoded Library for Chemical Assessment (BELKA) DNA-Encoded Library (DEL). Each structure includes at least one bound small molecule ligand, providing a robust basis for benchmarking model performance in accurately identifying potential binders to BELKA protein targets.
Introduction to the BELKA… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/BELKA-DEL-Experimental-BenchmarkData.GlycoShape
GlycoShape: A structural dataset of glycans
GlycoShape is a curated dataset of carbohydrate conformational ensembles derived from molecular dynamics simulations and structural analysis.
The dataset aims to provide standardized structural representations of glycans and glycosidic torsions to support computational glycobiology,
structural bioinformatics, and machine learning applications.
The dataset provides torsional information, conformational clusters, and representative… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/GlycoShape.
