datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
protein_backbone_cath_4.3
Dataset Description
https://github.com/facebookresearch/esm/tree/main/examples/inverse_folding#data-split
protein_backbone_cath_4.2
Dataset Description
https://github.com/jingraham/neurips19-graph-protein-design/tree/master/data
https://people.csail.mit.edu/ingraham/graph-protein-design/data/cath/
uniref-backboneref-OMG-Prot50-len_50_1024_entropy_3.5laurashin_Oracles_Does_the_Backbone_of_DeFi_Need_Fixing_-_Ep__537backboneref-testbackboneref
Documentation: BackboneRef Dataset
Overview
The BackboneRef dataset is a synthetic protein sequence resource developed as part of the Dayhoff Atlas. It was designed to enhance protein language model (PLM) training by introducing structurally informed diversity through de novo generated protein backbones and corresponding designed sequences. The dataset integrates structural novelty and plausibility to enable better generalization in sequence generation tasks.… See the full description on the dataset page: https://huggingface.co/datasets/fredzzp/backboneref.backbones
