datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
esmfold2-assets
LevinHarness/esmfold2-assets — internal asset cache
This private dataset is an internal staging cache of third-party runtime
assets required by the Levin Harness plugin(s) listed below. It is not an
official distribution: nothing here is published under this account's own
terms, and it is not affiliated with or endorsed by any upstream project.
Ownership and licensing
Every file remains the property of its upstream authors.
Each file keeps its upstream license… See the full description on the dataset page: https://huggingface.co/datasets/LevinHarness/esmfold2-assets.PDB-Monomeric-Structure-ESMFold2
PDB-Monomeric-Structure-ESMFold2
Monomeric, protein-only PDB structure dataset for minimum ESMFold2-style
training. Each row is one eligible single-chain biological assembly with a
canonical amino-acid sequence input and all-atom protein labels in atom37.
Labels
atom37_positions: residue x 37 x 3 coordinates, with zeros for missing atoms.
atom37_mask: residue x 37 resolved-atom mask.
aatype, residue_index, auth_seq_id, insertion_code, residue_name, ca_mask.… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/PDB-Monomeric-Structure-ESMFold2.esmfold1-fp8-wasmDeepLocMulti_ESMFold
DeepLocMulti Dataset with ESMFold Structural Sequence
Description: Protein localization encompasses the processes that establish and maintain proteins at specific locations.
Number of labels: 10
Problem Type: single_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
location: Specific location
Github
Simple, Efficient and Scalable Structure-aware Adapter… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepLocMulti_ESMFold.MetalIonBinding_ESMFold
MetalIonBinding Dataset with ESMFold Structural Sequence
Description: Metal-binding proteins are proteins or protein domains that chelate a metal ion.
Number of labels: 2
Problem Type: single_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
ss3_seq: DSSP 3 secondary structure sequence
esm3_structure_seq: ESM3 structure sequence encoded by VQ-VAE
Github… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/MetalIonBinding_ESMFold.GO_MF_ESMFold
GO-MF Dataset with ESMFold Structural Sequence
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF_ESMFold.GO_BP_ESMFold
GO-BP Dataset with ESMFold Structural Sequence
Description: Biological Process of Gene Ontology (GO) project.
Number of labels: 1943
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_BP_ESMFold.Thermostability_ESMFold
Thermostability Dataset with ESMFold Structural Sequence
Description: In materials science and molecular biology, thermostability is the ability of a substance to resist irreversible change in its chemical or physical structure, often by resisting decomposition or polymerization, at a high relative temperature.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/Thermostability_ESMFold.GO_CC_ESMFold
GO-CC Dataset with ESMFold Structural Sequence
Description: Cellular Component of Gene Ontology (GO) project.
Number of labels: 320
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_CC_ESMFold.esmfold2-assets
LevinHarness/esmfold2-assets — internal asset cache
This private dataset is an internal staging cache of third-party runtime
assets required by the Levin Harness plugin(s) listed below. It is not an
official distribution: nothing here is published under this account's own
terms, and it is not affiliated with or endorsed by any upstream project.
Ownership and licensing
Every file remains the property of its upstream authors.
Each file keeps its upstream license… See the full description on the dataset page: https://huggingface.co/datasets/sgetttt/esmfold2-assets.DeepLocBinary_ESMFold
DeepLocBinary Dataset with ESMFold Structural Sequence
Description: Protein localization encompasses the processes that establish and maintain proteins at specific locations.
Number of labels: 2
Problem Type: single_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
location: On the membrane or not
Github
Simple, Efficient and Scalable Structure-aware… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepLocBinary_ESMFold.DeepSol_ESMFold
DeepSol Dataset with ESMFold Structural Sequence
Description: Solubility is a fundamental protein property that has important connotations for therapeutics and use in diagnosis.
Number of labels: 2
Problem Type: single_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepSol_ESMFold.EC_ESMFold
EC Dataset with ESMFold Structural Sequence
Description: The Enzyme Commission number (EC number) is a numerical classification scheme for enzymes, based on the chemical reactions they catalyze.
Number of labels: 585
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/EC_ESMFold.ProtSolM_ESMFold_PDB
ProtSolM PDB Dataset (PDBSol)
Github
ProtSolM: Protein Solubility Prediction with Multi-modal Features
https://github.com/tyang816/ProtSolM
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@inproceedings{tan2024protsolm,
title={Protsolm: Protein solubility prediction with multi-modal features},
author={Tan… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/ProtSolM_ESMFold_PDB.DeepSol_ESMFold_PDB
DeepSol PDB Dataset
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2024ses-adapter,
title={Simple, Efficient, and Scalable Structure-Aware Adapter Boosts Protein… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepSol_ESMFold_PDB.ProtSolM_ESMFold
ProtSolM PDB Dataset (PDBSol)
Description: Solubility is a fundamental protein property that has important connotations for therapeutics and use in diagnosis.
Number of labels: 2
Problem Type: single_label_classification
Columns:
aa_seq: protein amino acid sequence
detail: meta information
Github
ProtSolM: Protein Solubility Prediction with Multi-modal Features
https://github.com/tyang816/ProtSolM
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/ProtSolM_ESMFold.GO_ESMFold_PDB
GO PDB Dataset
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2024ses-adapter,
title={Simple, Efficient, and Scalable Structure-Aware Adapter Boosts Protein Language… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_ESMFold_PDB.DeepSoluE_ESMFold
DeepSoluE Dataset with ESMFold Structural Sequence
Description: Solubility is a fundamental protein property that has important connotations for therapeutics and use in diagnosis.
Number of labels: 2
Problem Type: single_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepSoluE_ESMFold.Thermostability_ESMFold_PDB
Thermostability PDB Dataset
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2024ses-adapter,
title={Simple, Efficient, and Scalable Structure-Aware Adapter Boosts… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/Thermostability_ESMFold_PDB.DeepSoluE_ESMFold_PDB
DeepSolE PDB Dataset
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2024ses-adapter,
title={Simple, Efficient, and Scalable Structure-Aware Adapter Boosts Protein… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepSoluE_ESMFold_PDB.DeepLocMulti_ESMFold_PDB
DeepLocMulti PDB Dataset
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2024ses-adapter,
title={Simple, Efficient, and Scalable Structure-Aware Adapter Boosts… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepLocMulti_ESMFold_PDB.DeepLocBinary_ESMFold_PDB
DeepLocBinary PDB Dataset
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
Citation
Please cite our work if you use our dataset.
@article{tan2024ses-adapter,
title={Simple, Efficient, and Scalable Structure-Aware Adapter Boosts Protein Language Models},
author={Tan, Yang and Li, Mingchen and Zhou, Bingxin and Zhong, Bozitao and Zheng, Lirong and Tan, Pan and Zhou, Ziyi and… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepLocBinary_ESMFold_PDB.DeepET_Topt_ESMFold
DeepET_Topt Dataset
Description: protein optimum temperature.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
ss8_seq: DSSP 8 secondary structure sequence
foldseek_seq: foldseek 20 3di structural sequence
ss3_seq: DSSP 3 secondary structure sequence
Github
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepET_Topt_ESMFold.eSOL_ESMFold_PDB
Github
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2025venusfactory,
title={VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning},
author={Tan, Yang and Liu, Chen and Gao, Jingyuan and Wu, Banghao and Li, Mingchen and Wang, Ruilin and Zhang, Lingrong and… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/eSOL_ESMFold_PDB.MetalIonBinding_ESMFold_PDB
MetalIonBinding PDB Dataset
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2024ses-adapter,
title={Simple, Efficient, and Scalable Structure-Aware Adapter Boosts… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/MetalIonBinding_ESMFold_PDB.EC_ESMFold_PDB
EC PDB Dataset
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2024ses-adapter,
title={Simple, Efficient, and Scalable Structure-Aware Adapter Boosts Protein Language… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/EC_ESMFold_PDB.DeepET_Topt_ESMFold_PDB
Github
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2025venusfactory,
title={VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning},
author={Tan, Yang and Liu, Chen and Gao, Jingyuan and Wu, Banghao and Li, Mingchen and Wang, Ruilin and Zhang, Lingrong and… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/DeepET_Topt_ESMFold_PDB.proteinbase-esmfold2fast
proteinbase ESMFold2-Fast Structure Predictions
This dataset contains ESMFold2-Fast binder-target complex predictions.
Each prediction includes:
complex.cif
pae.npy
metrics.json
The top-level manifest.csv maps each prediction to the tar shard and stores pLDDT, pTM, iPTM, label metadata, and relative paths inside the archive.
Summary:
predictions included: 10890
predictions excluded during packaging: 90
seeds: 1, 2, 3, 4, 5
Download with:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/yk0/proteinbase-esmfold2fast.litscrape-esmfold2fast
litscrape ESMFold2-Fast Structure Predictions
This dataset contains ESMFold2-Fast binder-target complex predictions.
Each prediction includes:
complex.cif
pae.npy
metrics.json
The top-level manifest.csv maps each prediction to the tar shard and stores pLDDT, pTM, iPTM, label metadata, and relative paths inside the archive.
Summary:
predictions included: 20100
predictions excluded during packaging: 0
seeds: 1, 2, 3, 4, 5
Download with:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/yk0/litscrape-esmfold2fast.VenusVaccine_BacteriaBinary_ESMFold
