datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
malaria_CIDs_SIDs_SMILES_targets
Dataset was extracted from a dataset provided by NOVARTIS: Inhibition of Plasmodium falciparum W2 (drug-resistant) proliferation in erythrocyte-based infection assay https://pubchem.ncbi.nlm.nih.gov/bioassay/449704
Code of NNs https://pubchem.ncbi.nlm.nih.gov/bioassay/449704 utilising the resulting dataset.
human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_conciseThe human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_concised dataset is a part of the study "Leveraging 13C NMR spectroscopic data derived from SMILES to predict the functionality of small biomolecules by machine learning: a case study on human Dopamine D1 receptor antagonists "
https://doi.org/10.48550/arXiv.2501.14044
Dataset content: The Dataset has 59,567 rows of samples and 224 columns of features. Each feature corresponds to a range of the 13C NMR spectroscopy chemical… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_concise.CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_conciseThe CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_concise dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content: The Dataset has 20,309 rows of samples, 12 columns of features. Each feature corresponds to a range of the 1H NMR spectroscopy chemical shifts scale. Natural… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_concise.CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_extensiveThe CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_extensive dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content: The Dataset has 20,309 rows of samples, 103 columns of features. Each feature corresponds to a range of the 1H NMR spectroscopy chemical shifts scale.… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_extensive.CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_conciseThe CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_concise dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content: The Dataset has 20,309 rows of samples, 220 columns of features. Each feature corresponds to a range of the 13C NMR spectroscopy chemical shifts scale.… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_concise.CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_extensiveThe CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_extensive dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content:
The Dataset has 20,309 rows of samples, 1,928 columns of features. Each feature corresponds to a range of the 13C NMR spectroscopy chemical shifts scale.… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_extensive.human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_with_molecular! A note: To address the size limitations on Hugging Face, only 200 of the 59,609 rows were uploaded. The full dataset is available upon request for interested parties
The human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_with_molecular_features dataset is a part of the study " Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists"
https://doi.org/10.48550/arXiv.2506.01137
The… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_with_molecular.mke-novel-druglike-smiles
MKE Novel Drug-Like Molecules — AI-Generated SMILES Dataset
Commercial dataset available for purchase. Get access on our website →
Overview
This dataset contains 9,274 AI-generated, novel, drug-like small molecules, rigorously validated and filtered for pharmaceutical relevance. Every molecule in this dataset is:
✅ Chemically valid — RDKit-verified SMILES
✅ 100% novel — verified against 4,643,595 known compounds from MOSES, ZINC-250k, and ChEMBL
✅ Drug-like — QED… See the full description on the dataset page: https://huggingface.co/datasets/MKEChem/mke-novel-druglike-smiles.SMILES_Big_Data_Set
