datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PDB
PDB mmCIF Entry Index
The Protein Data Bank is the single global archive of experimentally-determined 3D structures of biological macromolecules, established in 1971 and now holding well over 230,000 entries. It stores atomic coordinates for proteins, nucleic acids, and their complexes determined by X-ray crystallography, cryo-EM, NMR, micro-electron diffraction, and integrative methods, along with the underlying experimental data (structure factors, EM maps, NMR restraints) and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/PDB.PDB
Data types
sequence: 355 datapoints
structure: 355 datapoints
pdb-dbpdbCDDB-PDB-Protein-50-512
CDDB–PDB Protein Structures, 50–512 Residues
Curated sequences, observed atomic coordinates, physical side-chain torsions,
observation masks, and separately filtered intrinsic-backbone labels for protein
generation and conditional modeling.
Training PDB cutoff: 31 December 2023, inclusive, using the entry's
initial public release date. The source snapshot was collected on
11 September 2026. These dates serve different purposes: historical entries use
their audited coordinates… See the full description on the dataset page: https://huggingface.co/datasets/alegendaryfish/CDDB-PDB-Protein-50-512.pdbbind_complexesA dataset to fine-tune language models on protein-ligand binding affinity and contact prediction.PDB-Monomeric-Structure-ESMFold2
PDB-Monomeric-Structure-ESMFold2
Monomeric, protein-only PDB structure dataset for minimum ESMFold2-style
training. Each row is one eligible single-chain biological assembly with a
canonical amino-acid sequence input and all-atom protein labels in atom37.
Labels
atom37_positions: residue x 37 x 3 coordinates, with zeros for missing atoms.
atom37_mask: residue x 37 resolved-atom mask.
aatype, residue_index, auth_seq_id, insertion_code, residue_name, ca_mask.… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/PDB-Monomeric-Structure-ESMFold2.PDB_primary_citation
primary_citation_with_pmcid.jsonl
This dataset links PDB protein structures with their corresponding primary ciatation text content.
Format
Each line is a JSON object:
{
"protein_name": "9HCG",
"structure_title": "Mouse mitoribosome large subunit assembly intermediate bound to NSUN4, MTERF4, and mt-RNAs",
"main_text": "..."
}
protein_structure_NER_independent_val_set
Overview
This data was used to evaluate the two models below to decide whether convergence was reached.
https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v2.1
https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v3.1
There are 20 different entity types in this dataset:
"bond_interaction", "chemical", "complex_assembly", "evidence", "experimental_method", "gene",
"mutant", "oligomeric_state", "protein", "protein_state", "protein_type"… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_independent_val_set.pdbookspdbbind_affinitiespdb_sequences
PDB Sequences
This dataset contains 780,163 protein sequences from the RCCB Protein Data Bank
Archive_CollaGNN_PDB_txtpdb_swissprotPDB-Chain-Complex-Benchmark-Rigor
PDB-Chain-Complex-Benchmark-Rigor
Rigor rebuild of Synthyra/PDB-Chain-Complex-Benchmark with split assignments
recomputed from the published chain and complex parquet artifacts.
Split Policy
Splits are assigned by connected components over exact sequence, sequence hash,
30% sequence cluster, structure cluster, source split component, same-PDB
asymmetric-unit membership, chain assembly membership, and biological assembly
co-membership from the complex rows. The… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/PDB-Chain-Complex-Benchmark-Rigor.PDB_Ribonanzanetpdbbind_complex_GB2022
To generate the dataset
Register for an account at https://www.pdbbind.org.cn/, confirm the validation email, then login and download
the Index files (1)
the general protein-ligand complexes (2)
the refined protein-ligand complexes (3)
Extract those files in pdbbind_complex_GB2022/data
Run the script pdbbind.py in a compute job on an MPI-enabled cluster (e.g., mpirun -n 64 pdbbind.py).
Output will be tar files in train/, val/ and test/ folders, following the split direction which is… See the full description on the dataset page: https://huggingface.co/datasets/djh992/pdbbind_complex_GB2022.pdbbind_refinedpdbbind_fullPDB-Single-Hard
PDB-Single-Hard: Precise Debugging Benchmarking — hard single-line bug subset
📄 Paper ·
💻 Code ·
🌐 Project page ·
🏆 Leaderboard
PDB-Single-Hard is the hard single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
Source datasets: BigCodeBench + LiveCodeBench
Sibling datasets: PDB-Single ·… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Hard.ecoli_pdb_benchmarkPDB-Single
PDB-Single: Precise Debugging Benchmarking — single-line bug subset
📄 Paper ·
💻 Code ·
🌐 Project page ·
🏆 Leaderboard
PDB-Single is the single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
Source datasets: BigCodeBench + LiveCodeBench
Sibling datasets: PDB-Single-Hard · PDB-Multi… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.pdb-datasets
Datasets for Cryo-EM synthetic training
Existing datasets:
scope: collection of PDBs categorized from scope
tim-rossman: collection of tim and rossman PDBs
Dataset Structure:
{dataset}/
- pdbs: all pdbs in the dataset
- dataframes: different types of datasets (how many families, which pdbs selected, etc.). All paths are relative
PDB-Chain-Complex-Benchmark
PDB-Chain-Complex-Benchmark
PDB-derived protein chain and biological assembly benchmark with strict
sequence, sequence-cluster, structure-cluster, and component-disjoint splits.
Configs
chains: one row per protein polymer chain instance.
complexes: one row per biological assembly with list-valued member chains.
Split Policy
Rows are split by connected components over exact sequence duplicate groups,
30% MMseqs2 sequence clusters, Foldseek… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/PDB-Chain-Complex-Benchmark.protein_chain_conformational_states
Schema description:
The manually curated dataset of open-closed monomers is included here as benchmarking_monomeric_open_closed_conformers.csv.
Column descriptions:
Schema description:
The manually curated dataset of open-closed monomers is included here as benchmarking_monomeric_open_closed_conformers.csv.
Column descriptions:
UNP_ACC | UniProt accession code
UNP_START | Start of UniProt sequence for given PDBe entries
UNP_END | End of UniProt sequence for given… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_chain_conformational_states.PDBenchpd_books_samplesPDB-Multi
PDB-Multi: Precise Debugging Benchmarking — multi-line bug subset (2–4 line blocks)
📄 Paper ·
💻 Code ·
🌐 Project page ·
🏆 Leaderboard
PDB-Multi is the multi-line bug subset (2–4 line blocks) of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
Source datasets: BigCodeBench + LiveCodeBench
Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi.PDB_snapshot_mmCIF_20250101reference_3d_validity_pdb
