CoolFace
Datasetpublic

ivanovaml/neuronalTTR_targetsTranscriptionActivators_C13NMR

The neuronalTTR_targetsTranscriptionActivators_13CNMR dataset is a part from the study “Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists” https://doi.org/10.48550/arXiv.2506.01137 A total of 3,041 unique small molecule samples are included in this dataset. The samples are classified by their TTR transcription activity, resulting in 1,093 activators and 1,948 non-activators. This information… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/neuronalTTR_targetsTranscriptionActivators_C13NMR.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes11downloads
Dataset Card

The neuronalTTRtargetsTranscriptionActivators13CNMR dataset is a part from the study “Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists” https://doi.org/10.48550/arXiv.2506.01137

A total of 3,041 unique small molecule samples are included in this dataset. The samples are classified by their TTR transcription activity, resulting in 1,093 activators and 1,948 non-activators. This information, along with the other molecular properties, is consistently defined across 211 columns, which are listed below: :

  • —Column 1: SMILES (Simplified Molecular Input Line Entry System) notations of the considered small biomolecules
  • —Columns 2 - 209 : Counts of the 13CNMR spectrosopy chemical shifts corresponding to the relevent small biomolecule. Names of the fatures corresponding to the range in which the picks have been counted.
  • —Columns 210: CID - PubChem identifier of the consdered compound
  • —Columns 211: target - The labels indicate whether the molecule is a TTR transcription activator: “1”, or not so: “0”.

The PubChem CIDs of the considered small biomolecules, their SMILES notations and corresponding labels regarding TTR transcription activation were provided by National Institutes of Health (NIH), PubChem AID 1117267 bioassay https://pubchem.ncbi.nlm.nih.gov/bioassay/1117267 "Luminescence-based cell-based primary high throughput screening assay to identify activators of Transthyretin (TTR) transcription", that contains 1,155 active out of 91,943 compounds. The inactive samples were reduced by merging the dataset presented above above with the dataset of PubChem AID 1996 bioassay https://pubchem.ncbi.nlm.nih.gov/bioassay/1996 "Aqueous Solubility from MLSMR Stock Solutions", on SMLES, keeping only the common compounds for both bioassays.