keural-nova
keural-nova-v1.2-sft
Keural Nova v1.2 — SFT dataset (PRIVATE)
The supervised fine-tuning mix used to train Keural Nova v1.2. Cleaned & balanced:
benchmark test-splits excluded, code AST-validated, identity de-contaminated,
AI-Hub / non-commercial rows removed. Each row carries accurate source_name and license fields.
Total 146,274 rows — train 144,812 / eval 1,462.
Format: {"messages":[{"role","content"},...], "source_name":..., "license":...}
Composition (by source)… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-v1.2-sft.keural-nova-tooluse
Keural Nova — tool-calling SFT slice (PRIVATE)
Tool-calling data used for Keural Nova v1.2. 17,337 rows: single-turn, multi-turn
(call -> tool result -> final answer), negative (no-call), and long-context up to 32k tokens.
Rendered as Qwen XML tool calls via ms-swift's native agent schema
(tool_call/tool roles + per-row tools JSON string).
tooluse_short.jsonl — 14,337 rows (<= ~3.6k tokens)
tooluse_long32k.jsonl — 3,000 rows (8k–32k tokens)
Sources / licenses:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-tooluse.keural-nova-identity
Keural Nova — identity SFT data (PRIVATE)
MKD-original. 200 rows (ko+en) teaching the assistant it is Keural, developed by MKD,
including denials of Qwen / GPT / Gemini / Claude / Llama etc. Facts grounded in https://mkd.kr.
Format: {"messages":[{"role":"user"...},{"role":"assistant"...}]}. Mix with heavy replay when
training (identity-only over-fits). Released Apache-2.0.
