datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
28mil_milestonepubchem-cid-smiles-title-inchikey-28M
Dataset Description
This dataset contains chemical molecular information in SMILES representation and other related metadata, extracted from the PubChem Compound Extras FTP directory.
Data Source
The SMILES data used to create this dataset can be found from the following PubChem FTP location:
PubChem Compound Extras
KamusOne-28M-Indonesian
KamusOne (Kamus-1) is a synthethic Indonesian language dataset, generated by Mixtral8x7B.
About
This dataset was generated by Mixtral 8x7B. For the procedure, Mixtral is instructed that it will act as an Indonesian language dictionary, a native Indonesian speaker, etc. and that it will explain the meaning of a series of Indonesian words. Hence, the name of the dataset ("Kamus", literally "dictionary"). Construction of the word list goes like this. First, we extracted word frequency… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/KamusOne-28M-Indonesian.
