datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
typhoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-sovereign-capability-dataset.nz-sovereign-sft
nz-sovereign-sft
Status: placeholder. There is no data here yet. Reserved for Step 5 of the
Sovereign Agentic Pipeline,
November 2026.
What this will be
The instruction-tuning set used to train
nz-sovereign-mistral-lora.
Target is roughly 1000 pairs of New Zealand material, built by hand rather than scraped,
with provenance recorded for every item.
Provenance rules this dataset inherits
Every number and every claim traces to a primary source with a… See the full description on the dataset page: https://huggingface.co/datasets/sentry-ai/nz-sovereign-sft.
