datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smiles-molecules-chembl
ChEMBL Molecule Generation Dataset
Dataset Description
ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties measured by some oracles.… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-chembl.ta-ESM2
taxonomy_aware_ESM2
This repository implements a Taxonomy-Aware Protein Function Prediction model. It synergizes the structural language understanding of ESM2 (Evolutionary Scale Modeling) with explicit phylogenetic lineage information.
smiles-molecules-moses
MOSES Molecule Generation Dataset
Dataset Description
Molecular Sets (MOSES) is a benchmark platform for distribution learning based molecule generation. Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization. It is processed from the ZINC Clean Leads dataset.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-moses.malaria_CIDs_SIDs_SMILES_targets
Dataset was extracted from a dataset provided by NOVARTIS: Inhibition of Plasmodium falciparum W2 (drug-resistant) proliferation in erythrocyte-based infection assay https://pubchem.ncbi.nlm.nih.gov/bioassay/449704
Code of NNs https://pubchem.ncbi.nlm.nih.gov/bioassay/449704 utilising the resulting dataset.
SELFormer-smilesCHOP_inhibitors_SMILES_1H_NMR_spectrosopy_conciseThe CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_concise dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content: The Dataset has 20,309 rows of samples, 12 columns of features. Each feature corresponds to a range of the 1H NMR spectroscopy chemical shifts scale. Natural… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_concise.CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_extensiveThe CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_extensive dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content: The Dataset has 20,309 rows of samples, 103 columns of features. Each feature corresponds to a range of the 1H NMR spectroscopy chemical shifts scale.… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_1H_NMR_spectrosopy_extensive.CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_conciseThe CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_concise dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content: The Dataset has 20,309 rows of samples, 220 columns of features. Each feature corresponds to a range of the 13C NMR spectroscopy chemical shifts scale.… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_concise.bioactives-naturals-smiles-molgen
Valid Bioactives and Natural Product SMILES
~2.7M valid SMILES built and curated from ChemBL34 (Zdrazil et al. 2023), COCONUTDB (Sorokina et al. 2021), and Supernatural3 (Gallo et al. 2023) dataset.
Curated by: gbyuvd
References
BibTeX
COCONUTDB
@article{sorokina2021coconut,
title={COCONUT online: Collection of Open Natural Products database},
author={Sorokina, Maria and Merseburger, Peter and Rajan, Kohulan and Yirik, Mehmet Aziz and Steinbeck… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/bioactives-naturals-smiles-molgen.CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_extensiveThe CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_extensive dataset is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
Dataset content:
The Dataset has 20,309 rows of samples, 1,928 columns of features. Each feature corresponds to a range of the 13C NMR spectroscopy chemical shifts scale.… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_SMILES_13C_NMR_spectrosopy_extensive.human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_with_molecular! A note: To address the size limitations on Hugging Face, only 200 of the 59,609 rows were uploaded. The full dataset is available upon request for interested parties
The human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_with_molecular_features dataset is a part of the study " Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists"
https://doi.org/10.48550/arXiv.2506.01137
The… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_with_molecular.mke-novel-druglike-smiles
MKE Novel Drug-Like Molecules — AI-Generated SMILES Dataset
Commercial dataset available for purchase. Get access on our website →
Overview
This dataset contains 9,274 AI-generated, novel, drug-like small molecules, rigorously validated and filtered for pharmaceutical relevance. Every molecule in this dataset is:
✅ Chemically valid — RDKit-verified SMILES
✅ 100% novel — verified against 4,643,595 known compounds from MOSES, ZINC-250k, and ChEMBL
✅ Drug-like — QED… See the full description on the dataset page: https://huggingface.co/datasets/MKEChem/mke-novel-druglike-smiles.SMILES_Big_Data_Set
