datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kamus-besar-bahasa-indonesiaKamusOne-28M-Indonesian
KamusOne (Kamus-1) is a synthethic Indonesian language dataset, generated by Mixtral8x7B.
About
This dataset was generated by Mixtral 8x7B. For the procedure, Mixtral is instructed that it will act as an Indonesian language dictionary, a native Indonesian speaker, etc. and that it will explain the meaning of a series of Indonesian words. Hence, the name of the dataset ("Kamus", literally "dictionary"). Construction of the word list goes like this. First, we extracted word frequency… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/KamusOne-28M-Indonesian.Roleplay-Indonesian
RolePlay-Indonesian
Roleplay-Indonesian Dataset is a dataset for roleplaying in the Indonesian language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Indonesian.indonesian_instruct_storiesA dataset of parallel translation-based instructions for Indonesian language as a target language.
Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license.
The template IDs are:
(1, 'Terjemahkanlah penggalan teks cerita anak berikut dari teks berbahasa Inggris ke teks dalam Bahasa Indonesia:', 'Terjemahan atau padanan teks tersebut dalam Bahasa Indonesia adalah:'),
(2, 'Terjemahkanlah penggalan teks cerita anak berikut dari teks… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/indonesian_instruct_stories.
