CoolFace
Datasetpublic

proteinglm/temperature_stability

Dataset Card for Temperature Stability Dataset Dataset Summary The accurate prediction of protein thermal stability has far-reaching implications in both academic and industrial spheres. This task primarily aims to predict a protein’s capacity to preserve its structural stability under a temperature condition of 65 degrees Celsius. Dataset Structure Data Instances For each instance, there is a string representing the protein sequence… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/temperature_stability.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes85downloads
Dataset Card

Dataset Card for Temperature Stability Dataset

Dataset Summary

The accurate prediction of protein thermal stability has far-reaching implications in both academic and industrial spheres. This task primarily aims to predict a protein’s capacity to preserve its structural stability under a temperature condition of 65 degrees Celsius.

Dataset Structure

Data Instances

For each instance, there is a string representing the protein sequence and an integer label indicating whether the protein can maintain its structural stability at a temperature of 65 degrees Celsius. See the temperature stability dataset viewer to explore more examples.

{'seq':'MEHVIDNFDNIDKCLKCGKPIKVVKLKYIKKKIENIPNSHLINFKYCSKCKRENVIENL'
'label':1}

The average for the seq and the label are provided below:

FeatureMean Count
seq300

Data Fields

  • —seq: a string containing the protein sequence
  • —label: an integer label indicating the structural stability of each sequence.

Data Splits

The temperature stability dataset has 3 splits: train, valid, and test. Below are the statistics of the dataset.

Dataset SplitNumber of Instances in Split
Train283,057
Valid62,973
Test73,205

Source Data

Initial Data Collection and Normalization

We adapted the dataset strategy from TemStaPro.

Licensing Information

The dataset is released under the Apache-2.0 License.

Citation

If you find our work useful, please consider citing the following paper:

@misc{chen2024xtrimopglm,
  title={xTrimoPGLM: unified 100B-scale pre-trained transformer for deciphering the language of protein},
  author={Chen, Bo and Cheng, Xingyi and Li, Pan and Geng, Yangli-ao and Gong, Jing and Li, Shen and Bei, Zhilei and Tan, Xu and Wang, Boyan and Zeng, Xin and others},
  year={2024},
  eprint={2401.06199},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  note={arXiv preprint arXiv:2401.06199}
}