t5gemma
Datasets
All datasets matching “t5gemma”t5gemma2-indonesia-instruct-v1
T5Gemma-2 Indonesian Instruct — Mono-Repo
Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia.
Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder,
setiap config = folder dan berisi split train + validation (80:20) di level percakapan.
Struktur (by fungsi)
t5gemma2-indonesia-instruct-v1/
├── README.md
├── manifest.json
├── chat_idx_map.json
├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.t5gemma2-indonesia-chat-formatted
T5Gemma-2 Indonesian Chat & QA Dataset
A high-quality Indonesian language multi-turn conversation and reading comprehension dataset, specifically formatted for instruction tuning of sequence-to-sequence (Seq2Seq) models like T5-Gemma / T5-Gemma-2.
Dataset Description
This dataset contains over 7,400 multi-turn conversations and document-based Q&A in Bahasa Indonesia. It covers diverse topics including everyday life, technology, general knowledge, and structured… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-chat-formatted.t5-gemma-2-multimodal-embeddingnemotron_math_aops_c4_t5gemmareasoning_v1_20m_t5gemmat5gemma2-indonesia-vision-formatted
T5Gemma-2 Indonesian Vision & Multimodal Dataset
A high-quality Indonesian language visual instruction tuning and preference alignment dataset, specifically formatted for multimodal sequence-to-sequence (Seq2Seq) models like T5-Gemma-2 Vision (e.g., google/t5gemma-2-270m-270m).
This dataset provides embedded image features (datasets.Image) along with multi-turn conversations in Bahasa Indonesia, facilitating instruction-following and preference alignment training on… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-vision-formatted.
