CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cmncomp /moleculestext1B<n<10B0 likes1.8k downloads1y agoHugging Face02chao1224 /MoleculeSTM Dataset Specifications for MoleculeSTM We provide the raw dataset (after preprocessing) at this Hugging Face link. Or you can download them by running python download.py. 1. Pretraining Dataset: PubChemSTM For PubChemSTM, please note that we can only release the chemical structure information. If you need the textual data, please follow our preprocessing scripts. 2. Downstream Datasets Please refer to the following for three downstream tasks: DrugBank_data for… See the full description on the dataset page: https://huggingface.co/datasets/chao1224/MoleculeSTM.text10K<n<100K11 likes909 downloads3y agoHugging Face03language-plus-molecules /mCLM_Pretrain_AllBlockstext10M<n<100M0 likes372 downloads10mo agoHugging Face04Fangyinfff /molecules Molecules This repository stores downloaded 3D molecular datasets collected into one gated Hugging Face dataset repo. Included datasets geom_drugs: GEOM-Drugs (~7GB including QM9, ~304K). Status: uploaded. Uploaded: yes. Cleaned locally: yes, staged size 39.8 GB. geom_qm9: GEOM-QM9 (included in GEOM, ~130K). Status: uploaded. Uploaded: yes. Cleaned locally: yes, staged size 148.0 B. spice: SPICE v2 (~7GB, ~19K molecules / 1.1M conformers). Status: uploaded. Uploaded:… See the full description on the dataset page: https://huggingface.co/datasets/Fangyinfff/molecules.other0 likes220 downloads5mo agoHugging Face05antoinebcx /smiles-molecules-chembl ChEMBL Molecule Generation Dataset Dataset Description ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs. Task Description For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties measured by some oracles.… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-chembl.text1M<n<10M3 likes194 downloads2y agoHugging Face06marianna13 /mdpi_molecules_jsons0 likes153 downloads3y agoHugging Face07language-plus-molecules /mCLM_Pretrain_Alltext10M<n<100M0 likes84 downloads10mo agoHugging Face08antoinebcx /smiles-molecules-moses MOSES Molecule Generation Dataset Dataset Description Molecular Sets (MOSES) is a benchmark platform for distribution learning based molecule generation. Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization. It is processed from the ZINC Clean Leads dataset. Task Description For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-moses.text1M<n<10M2 likes67 downloads2y agoHugging Face09language-plus-molecules /mCLM_Pretrain_100ktext10M<n<100M0 likes63 downloads10mo agoHugging Face10language-plus-molecules /LPM-24_traintext100K<n<1M5 likes58 downloads2y agoHugging Face11muhammadhasyim /equiv-dens-paper-molecules equiv-dens Paper Molecules Train/valid/test datasets for H2O, ethanol, ethanethiol, and resorcinol used in the equiv_dens_ml paper experiments. Format Geometry files are pickled NumPy dicts with positions and atom_numbers. Density label files are lists of (mol_dict, calc_dict) PySCF tuples. Usage ./scripts/download_hf_datasets.sh --group paper python run.py train @config/training/h2o_small_all_001.txt Security note Files contain… See the full description on the dataset page: https://huggingface.co/datasets/muhammadhasyim/equiv-dens-paper-molecules.other0 likes43 downloads4mo agoHugging Face12marianna13 /layout-moleculesimagen<1K0 likes29 downloads3y agoHugging Face13Alan123 /molecules_completedtext1M<n<10M0 likes27 downloads1y agoHugging Face14language-plus-molecules /mCLM_Pretrain_10ktext1M<n<10M0 likes27 downloads10mo agoHugging Face15language-plus-molecules /LPM-24_train-extratext1M<n<10M0 likes24 downloads3y agoHugging Face16language-plus-molecules /LPM-24_eval-captiontext10K<n<100K0 likes21 downloads2y agoHugging Face17colabfit /flexible_molecules_JCP2021 Cite this dataset Vassilev-Galindo, V., Fonseca, G., Poltavsky, I., and Tkatchenko, A. flexible molecules JCP2021. ColabFit, 2023. https://doi.org/10.60732/71f8031b This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_i23sbm1o45sj_0 Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/flexible_molecules_JCP2021.tabular100K<n<1M0 likes20 downloads10mo agoHugging Face18SamsungSAILMontreal /Conjugated-xTB_2M_moleculesConjugated-xTB dataset of 2M OLED molecules from the paper arxiv.org/abs/2502.14842. 'f_osc' is the oscillator strength (correlated with brightness) and should be maximized to obtain bright OLEDs. 'wavelength' is the absorption wavelength, >=1000nm corresponds to the short-wave infrared absorption range, which is crucial for biomedical imaging as tissues exhibit relatively low absorption and scattering in NIR, allowing for deeper penetration of light. This is good dataset for training a… See the full description on the dataset page: https://huggingface.co/datasets/SamsungSAILMontreal/Conjugated-xTB_2M_molecules.tabular1M<n<10M3 likes18 downloads2y agoHugging Face19tjetka-ingenix /molecules-embd-demotext100K<n<1M0 likes17 downloads1y agoHugging Face20language-plus-molecules /LPM-24_eval-molgentext10K<n<100K1 likes15 downloads2y agoHugging Face21nirajandhakal /PubChem-Molecules1K<n<10K0 likes15 downloads1y agoHugging Face22language-plus-molecules /mCLM_Pretrain_1ktext1M<n<10M1 likes9 downloads9mo agoHugging Face23french-open-data /tableau-des-molecules-onereuses-et-dispositifs-medicaux-implantables-de-la-liste-en-sus Tableau des molécules onéreuses et dispositifs médicaux implantables de la liste en sus [!NOTE] Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Tableau des molécules onéreuses et dispositifs médicaux implantables de la liste en sus qui est disponible à l'adresse https://www.data.gouv.fr/datasets/53ec0897a3a729638b9265a1 Description L'objet est de présenter la répartition de 2011 à la dernière année scellée : pour… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/tableau-des-molecules-onereuses-et-dispositifs-medicaux-implantables-de-la-liste-en-sus.0 likes3 downloads10mo agoHugging Face24faraazakhtar185 /molecules_smiles0 likes3 downloads6mo agoHugging Face25mohovision /metanetx_molecules_selfies_tokenized0 likes2 downloads2y agoHugging Face26jonblustein /chemical_moleculestabular10K<n<100K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.