datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oasst-javanese
Dataset Summary
We translated the OpenAssistant Conversations (OASST) dataset into Javanese using Meta's No Language Left Behind (NLLB) model.
Why Javanese?
Javanese is spoken by over 90 million people on the island of Java in Indonesia. While its prevalence is comparable to other widely spoken languages, such as Vietnamese and Turkish, its representation in current large language model (LLM) chatbots remains limited. By translating this dataset, we aim to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/richardcsuwandi/oasst-javanese.javanese-hotel-receptionist-qna
Dataset Card for Alpaca-Cleaned
Repository: https://huggingface.co/datasets/7out/javanese-hotel-receptionist-qna
Dataset Description
This synthetic dataset is designed for training and fine-tuning language models to handle customer service inquiries in a hotel setting using Javanese language. The data has been generated in the Alpaca format to assist in building models that can follow customer service-related instructions and generate appropriate responses. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/7out/javanese-hotel-receptionist-qna.
