datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moda-general-capability-rollouts
MODA General Capability Retention Rollouts
This dataset contains the raw model generations and evaluation results for the
MODA general-capability retention experiments. It covers 16 models, seven
benchmarks, 260,592 prompt records, and 2,605,920 stored generations.
The evaluation code is pinned to source commit
12ea99b2a57a354f2b7d6792f62a3d9313192fa7.
Evaluation protocol
Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ,
HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.typhoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-sovereign-capability-dataset.typhoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-sovereign-capability-dataset.nanochat-brevo-capability-data-10x
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Eleven
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-10x.nanochat-brevo-capability-data
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Six
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders are… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data.nanochat-brevo-capability-data-v2
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training uses project-planning language;
validation uses evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and record order. Per-world… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-v2.nanochat-brevo-capability-v4-49k-20260714
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training balances
project_plan, build_manifest language; validation uses held-out
evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-v4-49k-20260714.luanti-capability-eval
Luanti Capability Evaluation Dataset
Description
Luanti Capability Evaluation Dataset for Luanti (Minetest) expertise fine-tuning.
Dataset Information
Size: 60 entries
Format: Harmony format for LLM fine-tuning
Source: Luanti ContentDB package collection
Quality: Filtered and validated Luanti package metadata
Usage
from datasets import load_dataset
dataset = load_dataset("ToddLLM/luanti-capability-eval")
print(dataset)
Schema
Each… See the full description on the dataset page: https://huggingface.co/datasets/ToddLLM/luanti-capability-eval.luanti-capability-training
Luanti Capability Training Dataset
Description
Luanti Capability Training Dataset for Luanti (Minetest) expertise fine-tuning.
Dataset Information
Size: 600 entries
Format: Harmony format for LLM fine-tuning
Source: Luanti ContentDB package collection
Quality: Filtered and validated Luanti package metadata
Usage
from datasets import load_dataset
dataset = load_dataset("ToddLLM/luanti-capability-training")
print(dataset)
Schema
Each… See the full description on the dataset page: https://huggingface.co/datasets/ToddLLM/luanti-capability-training.
