CoolFace
19 results

SELFIES

HoangHa /selfies-ids-cleanedtext100M<n<1B0 likes739 downloads2y agoHugging Facehheiden /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B8 likes689 downloads9mo agoHugging FaceHoangHa /belka-selfies-idstext10M<n<100M0 likes484 downloads2y agoHugging FaceHoangHa /belka-selfies-train-cls-fttabular10M<n<100M0 likes466 downloads2y agoHugging FaceHoangHa /smiles-selfies-pretraintext100M<n<1B3 likes462 downloads2y agoHugging FaceBilsteen /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/Bilsteen/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B0 likes436 downloads1mo agoHugging Face