datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rosetta-Activations
Rosetta Activations
Updated: 2026-06-15 02:30 UTC
Contrastive activation extractions for 17 semantic concepts across 46 language models,
supporting cross-architecture mechanistic interpretability research.
Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs
Papers: forthcoming
Dataset Structure
Rosetta-Activations/
├── rcp_v1/ # Current extraction line — richest data (N≈2000)
│ └── {Model_Name}/
│ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.SAbDab_raw
All raw data from The Structural Antibody Database (SAbDab)
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the cache directory, e.g.
$ HF_HOME=${HOME}/.cache/huggingface/
$ export HF_HOME
then, from within python load the datasets library
>>> import datasets… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab_raw.MIP
Microbiome Immunity Project: Protein Universe
~200,000 predicted structures for diverse protein sequences from 1,003
representative genomes across the microbial tree of life and annotate
them functionally on a per-residue basis.
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MIP.MegaScale
Mega-scale experimental analysis of protein folding stability in biology and design
The full MegaScale dataset contains 1,841,285 thermodynamic folding stability measurements
using cDNA display proteolysis of natural and designed proteins. From these 776,298 high-quality folding
stabilities (dataset2) cover all single amino acid variants and selected double mutants of 331 natural
and 148 de novo designed protein domains 40–72 amino acids in length. Of these mutations, 607,839 have… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MegaScale.SAAINTDB
SAAINTDB
This dataset is a curated version of the SAAINT-DB converted into a format compatible with the Hugging Face Datasets for machine learning applications.
The dataset contains 21,400 antibody entries derived from 11,304 PDB structures, reflecting the available structures as of February 2026. Each entry corresponds to an antibody chain and is uniquely identified using the PDB_ID_chain field (PDB ID + chain ID).
Dataset Splits
The dataset was split at the PDB… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAAINTDB.SAbDab
ML Application Curated SAbDab
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the cache directory, e.g.
$ HF_HOME=${HOME}/.cache/huggingface/
$ export HF_HOME
then, from within python load the datasets library
>>> import datasets
Load model datasets
To load… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab.esa-rosetta-observations
ESA Rosetta Observations
Credit: NASA/ESA
Part of the Solar System Datasets and Planetary Science Datasets collections on Hugging Face.
Complete observation metadata catalog from the ESA Rosetta mission to Comet 67P/Churyumov-Gerasimenko — 8,214,033
observations across 15 instruments.
Dataset description
Rosetta was ESA's groundbreaking mission to Comet 67P/Churyumov-Gerasimenko. Launched in 2004, it became the first spacecraft to orbit a comet (August 2014) and… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/esa-rosetta-observations.FireProtDB2
Dataset Card for FireProtDB_2.0
Subsets of protein stability data for single-point mutants from FireProtDB, a comprehensive curated database.
Dataset Details
Subsets of different thermal data of single-point mutations in the FireProtDB database with train/validation/test splits:
ΔG, ΔΔG
Tm, ΔTm
Fitness
Stabilizing
Dataset Description
This dataset contains curated subsets of various thermal stability measurements derived from FireProtDB. Subsets… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/FireProtDB2.AfCycDesign
Dataset Card for AfCycDesign
Hallucinated scaffolds used by AfCycDesign for cyclic peptide design.
Dataset Details
Sets 7-16 of hallucinated peptide cif files and experimental CCDC structures.
Dataset Description
This dataset contains hallucinated cyclic peptide scaffold structures (in CIF format) generated using AfCycDesign, a deep learning approach built on AlphaFold2 for de novo design of cyclic peptides. The scaffolds span peptide lengths of 7–16 residues… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/AfCycDesign.FPbase
FPbase: The Fluorescent Protein Database
FPbase is a free, open-source, community-editable database of fluorescent proteins and their properties, aimed at aggregating structured, searchable information useful to the imaging community and FP developers. Visit fpbase.org for more.
This dataset updated on ,March 1st, 2026, collects FPbase fluorescent protein records (e.g., names, identifiers, sequences, and photophysical properties) for downstream analysis and modeling.… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/FPbase.NAKBOriginal Paper:
Lawson CL, Berman HM, Vallat B, Chen L, Zirbel C (2024) The Nucleic Acid Knowledgebase: a new portal for 3D structural information about nucleic acids. Nucleic Acids Research 52, D245-D254.
https://doi.org/10.1093/nar/gkad957
Nucleic Acid Knowledgebase (NAKB)
NAKB data set contains 21166 structures including Nucleic Acids, Protein, and Ligand Annotations, and determined 3D structures found in the Nucleic Acid Database (NDB) and the Protein Data Bank (PDB), including… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/NAKB.PTMint
PTMint
This dataset is derived from PTMint (https://ptmint.sjtu.edu.cn/), (Post Translational Modifications that are associated with Protein-Protein Interactions) that contains manually curated complete experimental evidence of the PTM effecting on protein-protein interactions in multiple organisms, including H. sapines, A. thaliana, C. elegans, D. melanogaster, S. cerevisiae and S. pombe.
This Hugging Face dataset repository provides PTMint-derived tables including a precomputed… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/PTMint.UTexasAptamer
UT Aptamer Dataset
This is a collection of 1480 aptamer sequences from the University of Texas Aptamer Database as of 2023. This dataset is split into three subsets (train, test, and validation) based on clustering by CD-HIT.
Clustering
Clustering was conducting using the CD-HIT: Cluster Database at High Identity with Tolerance web browser using a 40% sequence identity threshold and word size of 2. To update this dataset with new reported aptamers, splits can be… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/UTexasAptamer.IsItABarrel
IsItABarrel
This dataset contains 1,881,712 sequences collected from 600 different bacterial proteomes with sequences ranked by their likelihood of encoding a TMBB.
QuickStart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the HuggingFace datasets library. First, from the command line install the datasets library
$ pip install datasets
Optionally set the cache directory, e.g.
$ HF_HOME=${HOME}/.cache/huggingface/
$… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/IsItABarrel.AbAgym
AbAgym
AbAgym is a curated dataset of deep mutational scanning (DMS) measurements for antibody-antigen complexes. This Hugging Face version reorganizes the original AbAgym files into loadable dataset configurations using Apache Parquet, while preserving the original structure archive.
The original AbAgym repository describes the dataset as containing 68 DMS datasets on antibody-antigen complexes, approximately 324,000 non-redundant mutations, 36,541 non-redundant interface mutations… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/AbAgym.PRIDE_Crosslinking_Archive
PRIDE Crosslinking Archive
This dataset aggregates publicly available crosslinking mass spectrometry (XL-MS) datasets from the PRIDE repository.
Each dataset is curated and categorized by crosslinking reagent a link type (inter-chain vs intra-chain). For intra-chain links where the protein can be mapped to a
UniProt ID, each link is mapped onto the corresponding AlphaFold Database (AFDB) structure, and the Cα-Cα distance for the linked residue pair is reported.
The result is a… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/PRIDE_Crosslinking_Archive.dataset-rosetta-testCatPred-DB
CatPred-DB: Enzyme Kinetic Parameters Database
Paper: CatPred: A comprehensive framework for deep learning in vitro enzyme kinetic parameters
GitHub: https://github.com/maranasgroup/CatPred-DB
Dataset Description
CatPred-DB contains the benchmark datasets introduced alongside the CatPred deep learning framework for predicting in vitro enzyme kinetic parameters. The datasets cover three key kinetic parameters:
Parameter
Description
Datapoints
kcat
Turnover… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/CatPred-DB.2J-Protein-Couplings2J-Protein-Coupling Dataset
This data set was curated from the paper below accessed through the Biological Magnetic Resonance Data Bank (BMRB). There are a total of 3999 2J coupling taken from 5 different proteins and up to 10 different experiments. This dataset contains information regarding PDB ID, Sequence, 2J coupling data of 15N, 13C, and 1H. Data was curated and organized into this set of the five papers below, with the addition of the sequence taken from the Protein Data Bank.
Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/2J-Protein-Couplings.BELKA-DEL-Experimental-BenchmarkData
BELKA-DEL-Experimental-BenchmarkData
This dataset comprises a curated collection of PDB structures, designed as an experimental validation benchmark for models trained on the Big Encoded Library for Chemical Assessment (BELKA) DNA-Encoded Library (DEL). Each structure includes at least one bound small molecule ligand, providing a robust basis for benchmarking model performance in accurately identifying potential binders to BELKA protein targets.
Introduction to the BELKA… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/BELKA-DEL-Experimental-BenchmarkData.GlycoShape
GlycoShape: A structural dataset of glycans
GlycoShape is a curated dataset of carbohydrate conformational ensembles derived from molecular dynamics simulations and structural analysis.
The dataset aims to provide standardized structural representations of glycans and glycosidic torsions to support computational glycobiology,
structural bioinformatics, and machine learning applications.
The dataset provides torsional information, conformational clusters, and representative… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/GlycoShape.rosetta-trialsproteinbase-rosetta-metricsa1_code_rosetta_eval_636d
mlfoundations-dev/a1_code_rosetta_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
15.7
58.5
74.6
28.6
42.2
44.4
29.8
5.5
8.8
AIME24
Average Accuracy: 15.67% ± 0.82%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
20.00%
6
30
2
13.33%
4
30
3
13.33%
4
30
4
13.33%
4
30
5… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_rosetta_eval_636d.FoldDock
FoldDock
A collection of 219 heterodimers from dockground benchmark 4 dataset, and 1503 heterodimeric structures from a recent study (Green, A. G. et al. Nat. Commun. 12, 1–12 (2021)) (dubbed “marks”). The benchmark dataset was used to train an original model, that model was used to predict the structures in the second dataset.
This dataset is split into three subsets, each with two splits corresponding to the groups used in the source (marks, dockground).
Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/FoldDock.MAHOMES_II
MAHOMES II
MAHOMES-II (Metal Activity Heuristic of Metalloprotein and Enzymatic Sites-II) is a structure-based dataset for classifying protein-bound metal sites as enzymatic or non-enzymatic. ('Enzyme' column)
Dataset Splits
The train/test split follows a temporal split based on PDB deposition date. Structures that were deposited before 2018 were included in the training dataset, and structures deposited in 2018 or later went into the holdout test set. The train dataset… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MAHOMES_II.litscrape-rosetta-metrics
