CoolFace
Datasetpublic

hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC

PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.

sourceHugging Facecc0-1.0updated 9mo agoView on Hugging Face
8likes689downloads

hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC · main · files are served by the source, never re-hosted here