datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moleculesMoleculeSTM
Dataset Specifications for MoleculeSTM
We provide the raw dataset (after preprocessing) at this Hugging Face link. Or you can download them by running python download.py.
1. Pretraining Dataset: PubChemSTM
For PubChemSTM, please note that we can only release the chemical structure information. If you need the textual data, please follow our preprocessing scripts.
2. Downstream Datasets
Please refer to the following for three downstream tasks:
DrugBank_data for… See the full description on the dataset page: https://huggingface.co/datasets/chao1224/MoleculeSTM.mCLM_Pretrain_AllBlockssmiles-molecules-chembl
ChEMBL Molecule Generation Dataset
Dataset Description
ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties measured by some oracles.… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-chembl.mCLM_Pretrain_Allsmiles-molecules-moses
MOSES Molecule Generation Dataset
Dataset Description
Molecular Sets (MOSES) is a benchmark platform for distribution learning based molecule generation. Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization. It is processed from the ZINC Clean Leads dataset.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-moses.mCLM_Pretrain_100kLPM-24_trainmCLM_Pretrain_10kmolecules_completedLPM-24_train-extraLPM-24_eval-captionflexible_molecules_JCP2021
Cite this dataset Vassilev-Galindo, V., Fonseca, G., Poltavsky, I., and Tkatchenko, A. flexible molecules JCP2021. ColabFit, 2023. https://doi.org/10.60732/71f8031b
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_i23sbm1o45sj_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/flexible_molecules_JCP2021.molecules-embd-demoConjugated-xTB_2M_moleculesConjugated-xTB dataset of 2M OLED molecules from the paper arxiv.org/abs/2502.14842.
'f_osc' is the oscillator strength (correlated with brightness) and should be maximized to obtain bright OLEDs.
'wavelength' is the absorption wavelength, >=1000nm corresponds to the short-wave infrared absorption range, which is crucial for biomedical imaging as tissues exhibit relatively low absorption and scattering in NIR, allowing for deeper penetration of light.
This is good dataset for training a… See the full description on the dataset page: https://huggingface.co/datasets/SamsungSAILMontreal/Conjugated-xTB_2M_molecules.LPM-24_eval-molgenmCLM_Pretrain_1k
