datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
protein-secondary-structure-netsurfp
NetSurfP-3.0 Secondary-Structure Splits
This dataset repo contains NetSurfP-derived protein secondary-structure labels
converted for Protein-I-JEPA probe training and evaluation.
Source page: https://services.healthtech.dtu.dk/services/NetSurfP-3.0/5-Dataset.php
Profile: hhblits
Labels are Q3 per-residue labels:
H: helix
E: beta strand
C: coil/other
.: ignored residue for loss and accuracy
Splits
Split
Rows
JSONL
TSV
train
10348
train.jsonl
tsv/train.tsv… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/protein-secondary-structure-netsurfp.protein_stability_single_mutation
Protein Data Stability - Single Mutation
This repository contains data on the change in protein stability with a single mutation.
Attribution of Data Sources
Primary Source: Tsuboyama, K., Dauparas, J., Chen, J. et al. Mega-scale experimental analysis of protein folding stability in biology and design. Nature 620, 434–444 (2023). Link to the paper
Dataset Link: Zenodo Record
As to where the dataset comes from in this broader work, the relevant dataset (#3) is shown in… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/protein_stability_single_mutation.protein-structure-trust-benchmark
Protein-Structure Trust-Routing Benchmark (Boltz-2)
Leakage-controlled benchmarks for confidence-calibrated trust routing over a protein-structure
predictor: given a specialist model's confidence (Boltz-2 ipTM / pLDDT) for a target, decide whether to
trust the prediction or pay to verify it — and score that decision against experimentally-measured
correctness. Evaluation substrate for the report "When does an LLM trust a specialist model? A cost-aware
trust-routing audit"… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/protein-structure-trust-benchmark.protein_secondary_structure_from_PDBThis dataset contains 125,955 protein sequences, with protein PDB ID, length, the sequence (primary structure), as well as secondary structure as identified from experiment. The shortest protein is composed of only 11 amino acids, along with the longest one that features up to 19,350 amino acids. The standard deviation of the length is 855 amino acids.
The dataset further includes overall secondary sturctrure content, for all eight classes of secondary structure types.
The beta sheet content… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/protein_secondary_structure_from_PDB.protein_structure_pathogenicity_dataset
Protein Structure Pathogenicity Dataset
Dataset Description
This dataset contains protein structures and metadata for benign and pathogenic missense variants, designed for training machine learning models to predict variant pathogenicity using protein structural information.
Dataset Summary
The dataset includes:
Protein 3D structures predicted via ESMFold
Benign and pathogenic variants derived from the ProteinGym benchmark
Structural and sequence… See the full description on the dataset page: https://huggingface.co/datasets/CHGGM-Aachen/protein_structure_pathogenicity_dataset.swissprot-proteins
Uniprot SwissProt v. 2024_01
List of protein sequences and selected protein-level annotations for SwissProt v. 2024_01.
References:
UniProt: the Universal Protein Knowledgebase in 2025. The UniProt Consortium. https://doi.org/10.1093/nar/gkae1010
protein_structure_uncertainty_auditor_v0.2Protein Structure Uncertainty Auditor
GoalDetect when predicted protein structures are too uncertain for downstream use.
Model must output
uncertainty_flag (yes/no)
uncertainty_type
recommendation
This dataset tests whether models can audit structural confidence before use in:
drug design
docking
mutation mapping
function inference
Run scorer
python scorer.py --predictions predictions.jsonl --test_csv data/test.csv
africa-supply-utilization-accounts-2010-proteins-year
Supply Utilization Accounts (2010-) — Proteins/Year | Africa (FAOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-supply-utilization-accounts-2010-proteins-year.protein-structure-prediction-dashboard-datasetUniRef90-GPCR-ProteinsAll UniRef90 sequences of G-protein coupled receptors (GPCR) class proteins across all species. G-protein coupled receptors are evolutionarily related proteins and cell surface receptors that detect molecules outside the cell in Eukariotes.
Contains both confirmed and putative proteins.
proteins_scanprosite
