CoolFace
Datasetpublic

IDEA-AI4S/PubChemSFT

all_clean.json remove overlapped parts with ChEBI-20 test remove no description SMILES Format:{ SMILES <str>: [ ["Please describe the molecule", DESCRIPTION], ..., ] } Stats max tokens length: 6113 min tokens length: 20 mean tokens length: 191 median tokens length: 149 Total 326,689 single turn dialogue. Total 293,302 SMILES examples. Size: Train: 264,391 Valid: 33,072 Test: 32,987 conversation template 'conversation':{ [ "from": "human", "value": <QUERY>, #… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/PubChemSFT.

sourceHugging Facemitupdated 2y agoView on Hugging Face
7likes119downloads
filetest.pkl94.0 MBdownload
filetrain.pkl750.3 MBdownload
filevalid.pkl93.7 MBdownload

IDEA-AI4S/PubChemSFT · main · files are served by the source, never re-hosted here