AOYPSK/lao_pairs_final
🇱🇦 Lao SFT Pairs Final A cleaned and merged Lao-language instruction-tuning dataset for supervised fine-tuning (SFT) of large language models — specifically built to improve Lao language capability in models like Gemma 4. Dataset Summary Split File Examples Train lao_train_final.jsonl 57,088 Validation lao_val_final.jsonl 2,978 Total 60,066 Data Sources This dataset merges two sources: 1. Lao continuation corpus (32.5%) Real… See the full description on the dataset page: https://huggingface.co/datasets/AOYPSK/lao_pairs_final.
This repository belongs to AOYPSK on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
