datamol-io
safe-gpt
SAFE Molecules Dataset (v2)
A large-scale molecular dataset containing approximately 1.17 billion unique molecules, each represented with both canonical SMILES and SAFE (Sequential Attachment-based Fragment Embedding) strings.
This dataset is intended to support large-scale pretraining and evaluation of chemical language models, including generative, conditional, and structure-aware modeling tasks.
Note
This is version 2 of the SAFE dataset. The original v1 release contained… See the full description on the dataset page: https://huggingface.co/datasets/datamol-io/safe-gpt.safe-drugs
SAFE
Sequential Attachment-based Fragment Embedding (SAFE) is a novel molecular line notation that represents molecules as an unordered sequence of fragment blocks to improve molecule design using generative models.
This is the drugs dataset used for benchmarking.
Find the details and how to use at SAFE in the repo https://github.com/datamol-io/safe or the paper https://arxiv.org/pdf/2310.10773.pdf.
