AOYPSK/lao_pairs_final
🇱🇦 Lao SFT Pairs Final A cleaned and merged Lao-language instruction-tuning dataset for supervised fine-tuning (SFT) of large language models — specifically built to improve Lao language capability in models like Gemma 4. Dataset Summary Split File Examples Train lao_train_final.jsonl 57,088 Validation lao_val_final.jsonl 2,978 Total 60,066 Data Sources This dataset merges two sources: 1. Lao continuation corpus (32.5%) Real… See the full description on the dataset page: https://huggingface.co/datasets/AOYPSK/lao_pairs_final.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face