datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-protocol-posttrain
THE PROTOCOL — post-training data
Training and evaluation data for post-training a small LLM (Qwen3-4B) to obey
"THE PROTOCOL", a deliberately nonsensical 18-rule behavior spec from a
CAIDAS / JMU Würzburg take-home assignment. Rules trigger on surface features
of the user message (length, language, casing, digits, keywords, ...), fire
in arbitrary subsets, and collide under a precedence scheme; the protocol is
written in English but must be applied to input in any language.
All… See the full description on the dataset page: https://huggingface.co/datasets/laolaorkk/the-protocol-posttrain.lao_pairs_final
🇱🇦 Lao SFT Pairs Final
A cleaned and merged Lao-language instruction-tuning dataset for supervised fine-tuning (SFT) of large language models — specifically built to improve Lao language capability in models like Gemma 4.
Dataset Summary
Split
File
Examples
Train
lao_train_final.jsonl
57,088
Validation
lao_val_final.jsonl
2,978
Total
60,066
Data Sources
This dataset merges two sources:
1. Lao continuation corpus (32.5%)
Real… See the full description on the dataset page: https://huggingface.co/datasets/AOYPSK/lao_pairs_final.
