datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coconut-chembl34-selfies-mlm
Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked)
This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks.
The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.PubChem10M_SELFIESPubChem10M dataset by DeepChem encoded to SELFIES using group-selfies.
PubChem10M_SELFIES_TokenizedCustom cl100k tokenized version of PubChem10M_SELFIES.
SELFormer-selfiessabrlo-chem-selfies-training
Valid Bioactives and Natural Product SELFIES
~1M valid SELFIES with seq_len<=25 using FastChemTokenizerSelfies, built and curated from ChemBL34 (Zdrazil et al. 2023), COCONUTDB (Sorokina et al. 2021), and Supernatural3 (Gallo et al. 2023) dataset.
Processing
The original dataset was processed by filtering max_seq_len<=25 for training ChemMiniQ3-SAbRLo. The cleaned dataset was then split into 6 approximately equal chunks for training in limited compute.
Curated by: gbyuvd… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/sabrlo-chem-selfies-training.zinc15_40m_selfies40M molecular structures (represented in SELFIES format) randomly selected from ZINC15 database.
Dataset Card Author
Nianze TAO
