datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moleculesMoleculeSTM
Dataset Specifications for MoleculeSTM
We provide the raw dataset (after preprocessing) at this Hugging Face link. Or you can download them by running python download.py.
1. Pretraining Dataset: PubChemSTM
For PubChemSTM, please note that we can only release the chemical structure information. If you need the textual data, please follow our preprocessing scripts.
2. Downstream Datasets
Please refer to the following for three downstream tasks:
DrugBank_data for… See the full description on the dataset page: https://huggingface.co/datasets/chao1224/MoleculeSTM.mCLM_Pretrain_AllBlocksmolecules
Molecules
This repository stores downloaded 3D molecular datasets collected into one gated Hugging Face dataset repo.
Included datasets
geom_drugs: GEOM-Drugs (~7GB including QM9, ~304K). Status: uploaded. Uploaded: yes. Cleaned locally: yes, staged size 39.8 GB.
geom_qm9: GEOM-QM9 (included in GEOM, ~130K). Status: uploaded. Uploaded: yes. Cleaned locally: yes, staged size 148.0 B.
spice: SPICE v2 (~7GB, ~19K molecules / 1.1M conformers). Status: uploaded. Uploaded:… See the full description on the dataset page: https://huggingface.co/datasets/Fangyinfff/molecules.smiles-molecules-chembl
ChEMBL Molecule Generation Dataset
Dataset Description
ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties measured by some oracles.… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-chembl.mdpi_molecules_jsonsmCLM_Pretrain_Allsmiles-molecules-moses
MOSES Molecule Generation Dataset
Dataset Description
Molecular Sets (MOSES) is a benchmark platform for distribution learning based molecule generation. Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization. It is processed from the ZINC Clean Leads dataset.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-moses.mCLM_Pretrain_100kLPM-24_trainequiv-dens-paper-molecules
equiv-dens Paper Molecules
Train/valid/test datasets for H2O, ethanol, ethanethiol, and resorcinol used in the equiv_dens_ml paper experiments.
Format
Geometry files are pickled NumPy dicts with positions and atom_numbers. Density label files are lists of (mol_dict, calc_dict) PySCF tuples.
Usage
./scripts/download_hf_datasets.sh --group paper
python run.py train @config/training/h2o_small_all_001.txt
Security note
Files contain… See the full description on the dataset page: https://huggingface.co/datasets/muhammadhasyim/equiv-dens-paper-molecules.layout-moleculesmolecules_completedmCLM_Pretrain_10kLPM-24_train-extraLPM-24_eval-captionflexible_molecules_JCP2021
Cite this dataset Vassilev-Galindo, V., Fonseca, G., Poltavsky, I., and Tkatchenko, A. flexible molecules JCP2021. ColabFit, 2023. https://doi.org/10.60732/71f8031b
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_i23sbm1o45sj_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/flexible_molecules_JCP2021.Conjugated-xTB_2M_moleculesConjugated-xTB dataset of 2M OLED molecules from the paper arxiv.org/abs/2502.14842.
'f_osc' is the oscillator strength (correlated with brightness) and should be maximized to obtain bright OLEDs.
'wavelength' is the absorption wavelength, >=1000nm corresponds to the short-wave infrared absorption range, which is crucial for biomedical imaging as tissues exhibit relatively low absorption and scattering in NIR, allowing for deeper penetration of light.
This is good dataset for training a… See the full description on the dataset page: https://huggingface.co/datasets/SamsungSAILMontreal/Conjugated-xTB_2M_molecules.molecules-embd-demoLPM-24_eval-molgenPubChem-MoleculesmCLM_Pretrain_1ktableau-des-molecules-onereuses-et-dispositifs-medicaux-implantables-de-la-liste-en-sus
Tableau des molécules onéreuses et dispositifs médicaux implantables de la liste en sus
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Tableau des molécules onéreuses et dispositifs médicaux implantables de la liste en sus qui est disponible à l'adresse https://www.data.gouv.fr/datasets/53ec0897a3a729638b9265a1
Description
L'objet est de présenter la répartition de 2011 à la dernière année scellée :
pour… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/tableau-des-molecules-onereuses-et-dispositifs-medicaux-implantables-de-la-liste-en-sus.molecules_smilesmetanetx_molecules_selfies_tokenizedchemical_molecules
