CoolFace
Datasetpublic

hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC

PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.

sourceHugging Facecc0-1.0updated 9mo agoView on Hugging Face
8likes689downloads
5 commits on main
d4240af9mo ago

Update README.md

hheiden
cc465009mo ago

Update README.md

hheiden
853f5e09mo ago

Upload README.md with huggingface_hub

hheiden
1e4c8839mo ago

Upload folder using huggingface_hub

hheiden
301b8a79mo ago

initial commit

hheiden