datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t5gemma2-indonesia-chat-formatted
T5Gemma-2 Indonesian Chat & QA Dataset
A high-quality Indonesian language multi-turn conversation and reading comprehension dataset, specifically formatted for instruction tuning of sequence-to-sequence (Seq2Seq) models like T5-Gemma / T5-Gemma-2.
Dataset Description
This dataset contains over 7,400 multi-turn conversations and document-based Q&A in Bahasa Indonesia. It covers diverse topics including everyday life, technology, general knowledge, and structured… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-chat-formatted.llama3_additional_rr40k_non_delete_sft_chat_formatllama3_regular_balanced_sft_chat_formatllama31_sft_non_delete_300k_chat_formatllama31_no_additional_chat_formatllama31_no_additional_chat_format_add40kllama31_no_additional_chat_format_add76k10k_llama31_first_wrong_math_chat_formatcounsel-chat-formatted
