CoolFace
Modelpublic

mrothroc/molgpt-moses-selfies-mixlab

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes20downloads
Model Card

MolGPT (MOSES), reproduced in mixlab, SELFIES

The same small MolGPT reproduction as the SMILES model, retrained on SELFIES instead of SMILES. SELFIES is a string representation where every valid string decodes to a valid molecule, so validity becomes a property of the representation rather than something the model has to learn. Trained in mixlab on a single Apple-silicon Mac.

This is a model to learn from and tinker with, not a state-of-the-art result. It is also a small worked example of a useful idea: when a property is hard for a model to learn, changing the representation can be a better lever than a bigger model. mixlab makes that kind of experiment fast to run.

What it is

  • —6,342,656 parameters: 8 layers, 256 dim, 8 heads, causal GPT. Identical to the SMILES model except the tokenizer and sequence length. Exported as a native Hugging Face GPT2LMHeadModel.
  • —SELFIES-symbol tokenizer, vocab 30.
  • —Trained on the MOSES benchmark, one molecule per sequence, 33,000 steps (about 10 epochs), roughly 5 hours on an M1 Max.

What the representation change buys

Scored on 30,000 unconditional samples with the official moses.metrics, next to the SMILES model:

MetricSMILES modelthis SELFIES model
Valid0.9881.000
Novelty0.7830.859
FCD/Test2.830.43
Filters0.9960.976

Validity is 1.000 by construction, Novelty and distribution fit both improve, and the MOSES medicinal-chemistry Filters rate slips a little (0.996 to 0.976). SELFIES is not ours (Krenn et al. 2020); the interesting part is that a representation choice, at no extra compute, delivered guaranteed validity and better fidelity.

Load and generate

It loads with stock transformers as a GPT2LMHeadModel, no trust_remote_code. Framing tokens: BOS=1, EOS=2, PAD=0. The model emits SELFIES symbols; decode them to SMILES with the selfies library.

python
import torch, selfies as sf
from transformers import AutoModelForCausalLM
from tokenizers import Tokenizer
from huggingface_hub import hf_hub_download

repo = "mrothroc/molgpt-moses-selfies-mixlab"
model = AutoModelForCausalLM.from_pretrained(repo).eval()             # GPT2LMHeadModel
tok = Tokenizer.from_file(hf_hub_download(repo, "tokenizer.json"))
BOS, EOS, PAD = 1, 2, 0

out = model.generate(torch.full((8, 1), BOS), do_sample=True, temperature=1.0,
                     top_k=0, max_length=73, eos_token_id=EOS, pad_token_id=PAD)
for row in out.tolist():
    seq = [t for t in row[1:] if t not in (BOS, PAD)]
    if EOS in seq: seq = seq[:seq.index(EOS)]
    selfies = "".join(tok.id_to_token(t) for t in seq)
    print(sf.decoder(selfies))                                        # valid SMILES, by construction

Learn more

The full recipe, including the representation lesson and the two other ways to guarantee validity (a grammar-constrained decoder and a post-hoc verifier), is in the mixlab cookbook: https://github.com/mrothroc/mixlab-cookbook. Trainer: mixlab.

For the story behind it, read You Don't Code the Model, You Edit the JSON.

Citation

This is a reproduction, so please cite what it builds on:

  • —MolGPT: Bagal, V.; Aggarwal, R.; Vinod, P. K.; Priyakumar, U. D. "MolGPT: Molecular Generation Using a Transformer-Decoder Model." J. Chem. Inf. Model. 2022, 62 (9), 2064-2076. doi:10.1021/acs.jcim.1c00600
  • —MOSES: Polykovskiy, D. et al. "Molecular Sets (MOSES): A Benchmarking Platform for Molecular Generation Models." Front. Pharmacol. 2020, 11, 565644. doi:10.3389/fphar.2020.565644
  • —SELFIES: Krenn, M.; Häse, F.; Nigam, A.; Friederich, P.; Aspuru-Guzik, A. "Self-Referencing Embedded Strings (SELFIES): A 100% Robust Molecular String Representation." Mach. Learn.: Sci. Technol. 2020, 1 (4), 045024. doi:10.1088/2632-2153/aba947