CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01maykcaldas /smiles-transformers smiles-transformers dataset TODO: Add references to the datasets we curated dataset features name: text Molecule SMILES : string name: formula Molecular formula : string name: NumHDonors Number of hidrogen bond donors : int name: NumHAcceptors Number of hidrogen bond acceptors : int name: MolLogP Wildman-Crippen LogP : float name: NumHeteroatoms Number of hetero atoms: int name: RingCount Number of rings : int name: NumRotatableBonds Number of rotable… See the full description on the dataset page: https://huggingface.co/datasets/maykcaldas/smiles-transformers.tabular1B<n<10B22 likes2.9k downloads3y agoHugging Face02hheiden /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B8 likes496 downloads8mo agoHugging Face03tlemenestrel /Smiles2Dockhttps://arxiv.org/pdf/2406.05738 text10M<n<100M1 likes433 downloads1y agoHugging Face04kjappelbaum /chemnlp_iupac_smiles Dataset Card for "chemnlp_iupac_smiles" More Information needed text10M<n<100M8 likes415 downloads3y agoHugging Face05Bilsteen /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/Bilsteen/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B0 likes395 downloads1mo agoHugging Face06hypnopump /smiles_transformer_shuffledtabular100M<n<1B0 likes318 downloads1y agoHugging Face07liyuesen /zinc20_smiles Dataset Card for "zinc20_smiles" More Information needed text1B<n<10B0 likes317 downloads1y agoHugging Face08HoangHa /smiles-selfies-pretraintext100M<n<1B3 likes300 downloads2y agoHugging Face09antoinebcx /smiles-molecules-chembl ChEMBL Molecule Generation Dataset Dataset Description ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs. Task Description For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties measured by some oracles.… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-chembl.text1M<n<10M3 likes196 downloads2y agoHugging Face10lovingscience /chembl-smiles-curated Dataset Card for Curated ChEMBL SMILES Dataset (lovingscience/chembl-smiles-curated) This dataset is a reproducibly cleaned, deduplicated, and split ChEMBL SMILES dataset built for training molecular generative models and chemical space mapping tools (such as Imolgen and Navifold). Dataset Pipeline Provenance & Configuration ChEMBL Release Version: 37 RDKit Version: 2026.03.6 Deduplication & Split Seed: 42 Split Ratios: Train (80.0%) / Validation (10.0%) / Test… See the full description on the dataset page: https://huggingface.co/datasets/lovingscience/chembl-smiles-curated.tabular1M<n<10M0 likes153 downloads8d agoHugging Face11th-laurel /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/th-laurel/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B0 likes147 downloads6mo agoHugging Face12jablonkagroup /ord_rxn_smiles_procedure Dataset Details Dataset Description The open reaction database is a database of chemical reactions and their conditions Curated by: License: CC BY SA 4.0 Dataset Sources original data source Citation BibTeX: @article{Kearnes_2021, doi = {10.1021/jacs.1c09820}, url = {https://doi.org/10.1021%2Fjacs.1c09820}, year = 2021, month = {nov}, publisher = {American Chemical Society ({ACS})}, volume = {143}, number = {45}, pages =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/ord_rxn_smiles_procedure.text100K<n<1M0 likes143 downloads1y agoHugging Face13IDEA-AI4S /RCR_RP_57K_SMILES-MMChatReaction Condition Prediction Dataset (Reagent Prediction) molecule representation format: 1D SMILES will further encode into 2D graph features For detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193 text10K<n<100K0 likes133 downloads2y agoHugging Face14suku9 /enamine_smiles_datasettext100M<n<1B0 likes133 downloads2y agoHugging Face15lovingscience /chembl-v37-smiles-curated Dataset Card for Curated ChEMBL SMILES Dataset (lovingscience/chembl-smiles-curated) This dataset is a reproducibly cleaned, deduplicated, and split ChEMBL SMILES dataset built for training molecular generative models and chemical space mapping tools (such as Imolgen and Navifold). Dataset Pipeline Provenance & Configuration ChEMBL Release Version: 37 RDKit Version: 2026.03.6 Deduplication & Split Seed: 42 Split Ratios: Train (80.0%) / Validation (10.0%) / Test… See the full description on the dataset page: https://huggingface.co/datasets/lovingscience/chembl-v37-smiles-curated.tabular1M<n<10M0 likes126 downloads8d agoHugging Face16eth-sri /smiles-eval SMILES eval This is a dataset that measures LLM capabilities at generating SMILES chemical molecule representations from natural language descriptions. It was generated by prompting Gemini 2.5 Pro for molecule description and SMILES pairs, and filtering for a) molecules valid according to rdkit and b) reliably regenerated by Gemini and c) removing duplicates using fuzzy matching. Difficulty is estimated by how reliable Gemini 2.5 Pro is at generating the molecule. This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/smiles-eval.textn<1K0 likes120 downloads1y agoHugging Face17jablonkagroup /pubchem-smiles-molecular-formulatext10M<n<100M4 likes115 downloads1y agoHugging Face18fabikru /chembl-2025-randomized-smiles-cleaned-rdkit-descriptorstabular1M<n<10M2 likes109 downloads1y agoHugging Face19Smilesjs /ta-ESM2 taxonomy_aware_ESM2 This repository implements a Taxonomy-Aware Protein Function Prediction model. It synergizes the structural language understanding of ESM2 (Evolutionary Scale Modeling) with explicit phylogenetic lineage information. text100K<n<1M0 likes106 downloads8mo agoHugging Face20mllab /smiles-2025text1K<n<10K0 likes101 downloads1y agoHugging Face21introvoyz041 /enamine_smiles_datasettext100M<n<1B0 likes98 downloads6mo agoHugging Face22IDEA-AI4S /MolInst_RS_125K_SMILES-MMChattext100K<n<1M0 likes96 downloads2y agoHugging Face23IDEA-AI4S /RCR_SP_70K_SMILES-MMChatReaction Condition Prediction Dataset (Solvent Prediction) molecule representation format: 1D SMILES will further encode into 2D graph features For Detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193 text10K<n<100K0 likes93 downloads2y agoHugging Face24IDEA-AI4S /MolInst_FS_125K_SMILES-MMChattext100K<n<1M0 likes92 downloads2y agoHugging Face25IDEA-AI4S /SMol_RS_Filtered_825K_SMILES-MMChatRetrosynthesis Prediction Dataset (derived from SMolInstruct) molecule representation format: 1D SMILES will further encode into 2D graph features We filtered out overlapping samples from the original train-split (test-set: MolInstruct-Retrosynthesis Prediction) We only include single-step retrosynthesis prediction. For Detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193 text100K<n<1M0 likes92 downloads2y agoHugging Face26IDEA-AI4S /RCR_CP_10K_SMILES-MMChatReaction Condition Prediction Dataset (Catalyst Prediction) molecule representation format: 1D SMILES will further encode into 2D graph features Detail refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193 text10K<n<100K1 likes87 downloads2y agoHugging Face27IDEA-AI4S /MolInst_RS_125K_Scaffold_SMILES-MMChatRetrosynthesis Prediction Dataset (derived from MolInstruct) molecule representation format: 1D SMILES will further encode into 2D graph features We use scaffold splitting to reconstruct the train-split. We use SMolInstruct RS train split as the sample pool. We only include single-step retrosynthesis prediction. For Detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193 text100K<n<1M1 likes80 downloads2y agoHugging Face28fabikru /pubchem_and_chembl-2025-randomized-smiles-cleanedtext100M<n<1B0 likes78 downloads2y agoHugging Face29keanec27 /Drug_Protein_Interactions_Smilestext10K<n<100K1 likes77 downloads2y agoHugging Face30IDEA-AI4S /SMol_FS_Filtered_875K_SMILES-MMChatForward Reaction Prediction Dataset (derived from SMolInstruct) molecule representation format: 1D SMILES will further encode into 2D graph features We filtered out overlapping samples from original train-split (test-set: MolInstruct-Forward Reaction Prediction) For Detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193 text100K<n<1M0 likes76 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.