datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
keural-SFT-chatml-ko-v1
Keural SFT ChatML (Korean) v1
한국어 SFT(Supervised Fine-Tuning)용 통합 데이터셋입니다. 공개 한국어 instruction/대화 데이터셋 8종을 수집하여 정제 → 품질 필터링 → 안전성 필터링 → 중복 제거(exact + MinHash near-dup) → ChatML 포맷팅 → 토크나이즈 검증을 거친 결과물입니다.
총 샘플 수: 710,278
총 토큰 수: 약 1.9억 (190,724,939 tokens, keural tokenizer 기준)
포맷: ChatML (<|im_start|>role ... <|im_end|>)
최대 길이: 8,192 tokens (초과 시 truncate)
생성일: 2026-07-10
데이터 구조
각 샤드는 JSONL 형식이며, 레코드 스키마는 다음과 같습니다:
{"text":… See the full description on the dataset page: https://huggingface.co/datasets/mkd-jueon/keural-SFT-chatml-ko-v1.keural-v2-dataset
Keural v2 — MoE Fine-Tuning Dataset
상태: 비공개 (private) — 공개 배포 대상 아님
Keural MoE Pro v2 모델 파인튜닝을 위해 8개 카테고리(A~H)로 구성된 SFT 학습 데이터셋입니다. 자세한 수집·처리 기준은 mkd-minju/Keural-MoE-Pro-v2 GitHub 저장소의 docs/V2-DATASET-PREP.md 계획 문서를 따릅니다.
카테고리 구성
파일
카테고리
목표 건수
언어
주요 출처
라이선스
A_korean_conversation.jsonl
한국어 대화/지침
50,000
ko
mkd-chanwoo/keural-conversation-chatml-ko, mkd-chanwoo/keural-rag-chatml-ko… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-dataset.keural-conversation-chatml-ko
keural-conversation-chatml-ko
mkd-chanwoo/keural-conversation-ko 데이터셋을 SFT 학습용 ChatML 포맷으로 전처리한 한국어 일상대화 데이터셋입니다.
9개 주제의 일상 캐주얼 대화를 Gemma-4-26B 모델이 생성한 synthetic 멀티턴 대화로 구성되어 있습니다.
데이터셋 개요
항목
값
총 샘플 수
133,339
총 토큰 수
약 19M
평균 토큰 / 샘플
146.0
중간값 토큰
145
최대 토큰
805
최소 토큰
95
언어
한국어 (100%)
포맷
ChatML
라이선스
CC BY 4.0
포맷
ChatML 형식의 멀티턴 대화입니다.
<|im_start|>user
{발화 1}
<|im_end|>
<|im_start|>assistant
{발화 2}
<|im_end|>
<|im_start|>user
{발화… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-conversation-chatml-ko.keural-nova-v1.2-sft
Keural Nova v1.2 — SFT dataset (PRIVATE)
The supervised fine-tuning mix used to train Keural Nova v1.2. Cleaned & balanced:
benchmark test-splits excluded, code AST-validated, identity de-contaminated,
AI-Hub / non-commercial rows removed. Each row carries accurate source_name and license fields.
Total 146,274 rows — train 144,812 / eval 1,462.
Format: {"messages":[{"role","content"},...], "source_name":..., "license":...}
Composition (by source)… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-v1.2-sft.keural-DPO
keural-DPO
한국어/영어 DPO(Direct Preference Optimization) 학습용 통합 데이터셋.
6개 소스 데이터셋을 ChatML 형식으로 정규화하여 통합한 것으로, 총 443,314 샘플을 포함합니다.
데이터셋 통계
Split
소스
필터 제거
업로드
언어
라이선스
ultrafeedback_binarized
61,054
424
60,630
EN
MIT
multifaceted_collection_dpo
65,139
1,045
64,094
EN/KO
Apache-2.0
hh_rlhf
160,608
739
159,869
EN
MIT
orca_dpo_pairs_ko
12,727
2
12,725
KO
—
aihub_71748
29,676
0
29,676
KO
AI Hub
aihub_71760
116,320
0
116,320
KO
AI Hub
합계
445,524
2,210
443,314… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-DPO.keural-cortex-8b-sft
Keural-Cortex-8B SFT dataset
The supervised fine-tuning set used to train Keural-Cortex-8B, a Korean-first
bilingual model with a 64K context window, tool calling, and hybrid
thinking/non-thinking modes.
1,568,649 rows · 1.87B estimated tokens · 1.80B real Qwen3 tokens · 73.7% Korean
Three files:
file
rows
what it is
train.jsonl
1,564,042
main set, all rows under 32,768 tokens
train_long64k.jsonl
4,607
the 32K–64K band, kept separate because it needs a different… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-cortex-8b-sft.keural-v2-tool-calling
Tool & Function Calling (Area 2) — Korean SFT Dataset Prep
상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님
출처
원본: glaiveai/glaive-function-calling-v2
커밋 해시: e7f4b6456019f5d8bcb991ef0dd67d8ff23221ac
라이선스: Apache-2.0 (원본 태그, README 본문 없어 대조 문구 없음)
생성 출처: 미확인 — GPT-4/Claude 등 프론티어 모델 사용 가능성 있음 (원본 데이터셋 카드에 명시 없음)
언어: 영어 (지침서 §1.2 정책에 따라 번역 없이 영어 그대로 사용)
처리 과정
원본 112,960건 다운로드
chat 필드 기준 완전 중복 23,790건(21%) 발견 및 제거 → 유니크 89,170건
유니크 풀에서 seed=42로 50… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-tool-calling.keural-v2-cot-reasoning
Reasoning / Chain-of-Thought (Area 4) — Korean SFT Dataset Prep
상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님
출처
원본: nvidia/OpenMathReasoning (cot split)
커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df
라이선스: CC-BY-4.0 (태그와 본문 일치, "License/Terms of Use: cc-by-4.0")
생성 모델: DeepSeek-R1(샘플 중 다수), QwQ-32B — 둘 다 오픈 웨이트 모델, 독점 모델 ToS 리스크 없음
언어: 영어 (지침서 §1.4 정책에 따라 번역 없이 영어 그대로 사용)
출처 구성 (problem_source)
문제(질문) 출처는 대부분 AoPS(Art of Problem Solving) 포럼… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning.keural-v2-self-verification
Self-Verification Dataset
Status: generation in progress. 4,477 / 50,000 target rows (~9%), growing. Being generated in parallel across multiple environments/models — see Generation models mix below. This card describes the file as of this snapshot; row count and model mix will change on re-upload.
File: 05_self_verification_generated.jsonl (one JSON object per line).
What this is
Synthetic "wrong draft → self-critique → corrected solution" traces, for training a… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-self-verification.keural-v2-orchestration
Agent Orchestration Dataset
Status: generation complete. 50,000 / 50,000 target rows. Not yet processed — see Pipeline stage before using this for training.
File: agent_orchestration_final_50000.jsonl (597 MB, one JSON object per line, 50,000 lines).
What this is
Synthetic multi-agent conversation traces for training a model to act as an orchestrator: decompose a user task, delegate subtasks to named sub-agents, then synthesize their independent responses into one… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-orchestration.keural-DPO-raw
keural-DPO-raw
keural-DPO의 전처리 전 normalized 원본 데이터.
필터링 없이 원본 그대로 보존하여 재현성을 위해 제공합니다.
ChatML 변환 및 필터링이 적용된 학습용 데이터는 mkd-chanwoo/keural-DPO 를 사용하세요.
데이터셋 통계
Split
행 수
언어
라이선스
원본
ultrafeedback
61,054
EN
MIT
HuggingFaceH4/ultrafeedback_binarized
multifaceted_dpo
65,139
EN/KO
Apache-2.0
kaist-ai/Multifaceted-Collection-DPO
hh_rlhf
160,608
EN
MIT
Anthropic/hh-rlhf
orca_dpo_ko
12,727
KO
—
Ja-ck/Orca-DPO-Pairs-KO
aihub_71748
29,676
KO
AI Hub
AI Hub 71748… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-DPO-raw.keural-v2-cot-reasoning-v2
Reasoning / Chain-of-Thought (Area 4, v2) — Korean SFT Dataset Prep
상태: 비공개 스테이징(private) — §3 처리(1~8번, 최종 인코딩 포함) 전부 완료. 제2자 감사 전, 공개 배포 대상 아님.
이 v2는 §3 처리를 새로 검증하며 진행한 최종 버전입니다(2026-08-10). v1(원본 problem/generated_solution 필드 그대로)과 달리, DeepSeek-V4-Flash-0731 학습용 최종 텍스트(text 필드)로 인코딩까지 완료됐습니다.
출처
원본: nvidia/OpenMathReasoning (cot split)
커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df
라이선스: CC-BY-4.0 (태그와 본문 일치)
생성 모델: DeepSeek-R1(다수), QwQ-32B — 둘 다 오픈 웨이트 모델
언어:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning-v2.keural-datasets-sampleskeural-v2-tool-calling-v2
Tool & Function Calling (Area 2, v2) — Korean SFT Dataset Prep
상태: 비공개 스테이징(private) — §3 처리(1~6번, 스키마 정규화) 완료, §3-7(최종 텍스트 인코딩)만 보류. 제2자 감사 전, 공개 배포 대상 아님.
이 v2는 §3 처리를 새로 검증하며 발견한 오류를 수정한 버전입니다(2026-08-10). v1(원본 chat 텍스트 그대로)과 달리, 이 저장소엔 korean_sft_schema.md 통합 구조(source/license/lang/category/conversations:[{role,content,reasoning_content,tool_calls}])로 정규화된 데이터가 들어있습니다.
출처
원본: glaiveai/glaive-function-calling-v2
커밋 해시: e7f4b6456019f5d8bcb991ef0dd67d8ff23221ac… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-tool-calling-v2.keural-nova-tooluse
Keural Nova — tool-calling SFT slice (PRIVATE)
Tool-calling data used for Keural Nova v1.2. 17,337 rows: single-turn, multi-turn
(call -> tool result -> final answer), negative (no-call), and long-context up to 32k tokens.
Rendered as Qwen XML tool calls via ms-swift's native agent schema
(tool_call/tool roles + per-row tools JSON string).
tooluse_short.jsonl — 14,337 rows (<= ~3.6k tokens)
tooluse_long32k.jsonl — 3,000 rows (8k–32k tokens)
Sources / licenses:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-tooluse.keural-nova-identity
Keural Nova — identity SFT data (PRIVATE)
MKD-original. 200 rows (ko+en) teaching the assistant it is Keural, developed by MKD,
including denials of Qwen / GPT / Gemini / Claude / Llama etc. Facts grounded in https://mkd.kr.
Format: {"messages":[{"role":"user"...},{"role":"assistant"...}]}. Mix with heavy replay when
training (identity-only over-fits). Released Apache-2.0.
