datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
keural-datasets
Keural Pretraining Datasets (Stage 2)
Stage 2 final production corpus for training the Keural Korean LLM.
Quality-filtered, deduplicated, and domain-balanced across 4 domains.
Summary
Metric
Value
Total processed documents (post-filter)
757,710,609
Dedup removed (Stage 2)
93,919,634
Final documents
663,790,975
Total tokens
~522B
Domains
English, Korean, Code, Science
Source datasets
43
Format
Parquet (snappy compressed, sharded)
Upload… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-datasets.Keural-MoE-14B-stage1-Datasetkeural-SFT-chatml-ko-v1
Keural SFT ChatML (Korean) v1
한국어 SFT(Supervised Fine-Tuning)용 통합 데이터셋입니다. 공개 한국어 instruction/대화 데이터셋 8종을 수집하여 정제 → 품질 필터링 → 안전성 필터링 → 중복 제거(exact + MinHash near-dup) → ChatML 포맷팅 → 토크나이즈 검증을 거친 결과물입니다.
총 샘플 수: 710,278
총 토큰 수: 약 1.9억 (190,724,939 tokens, keural tokenizer 기준)
포맷: ChatML (<|im_start|>role ... <|im_end|>)
최대 길이: 8,192 tokens (초과 시 truncate)
생성일: 2026-07-10
데이터 구조
각 샤드는 JSONL 형식이며, 레코드 스키마는 다음과 같습니다:
{"text":… See the full description on the dataset page: https://huggingface.co/datasets/mkd-jueon/keural-SFT-chatml-ko-v1.keural-v2-dataset
Keural v2 — MoE Fine-Tuning Dataset
상태: 비공개 (private) — 공개 배포 대상 아님
Keural MoE Pro v2 모델 파인튜닝을 위해 8개 카테고리(A~H)로 구성된 SFT 학습 데이터셋입니다. 자세한 수집·처리 기준은 mkd-minju/Keural-MoE-Pro-v2 GitHub 저장소의 docs/V2-DATASET-PREP.md 계획 문서를 따릅니다.
카테고리 구성
파일
카테고리
목표 건수
언어
주요 출처
라이선스
A_korean_conversation.jsonl
한국어 대화/지침
50,000
ko
mkd-chanwoo/keural-conversation-chatml-ko, mkd-chanwoo/keural-rag-chatml-ko… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-dataset.keural-conversation-chatml-ko
keural-conversation-chatml-ko
mkd-chanwoo/keural-conversation-ko 데이터셋을 SFT 학습용 ChatML 포맷으로 전처리한 한국어 일상대화 데이터셋입니다.
9개 주제의 일상 캐주얼 대화를 Gemma-4-26B 모델이 생성한 synthetic 멀티턴 대화로 구성되어 있습니다.
데이터셋 개요
항목
값
총 샘플 수
133,339
총 토큰 수
약 19M
평균 토큰 / 샘플
146.0
중간값 토큰
145
최대 토큰
805
최소 토큰
95
언어
한국어 (100%)
포맷
ChatML
라이선스
CC BY 4.0
포맷
ChatML 형식의 멀티턴 대화입니다.
<|im_start|>user
{발화 1}
<|im_end|>
<|im_start|>assistant
{발화 2}
<|im_end|>
<|im_start|>user
{발화… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-conversation-chatml-ko.keural-sft-mix
Keural SFT Mix (sampled)
A language/domain-balanced sample of mkd-chanwoo/keural-datasets,
drawn for SFT experimentation. Each category is kept as its own config (mirroring the source layout),
so the different category schemas never have to be merged.
Composition
Config
Rows
Size
Target mix
korean
1,238,411
~6.3 GB
40%
english
774,007
~2.1 GB
25%
code
619,205
~2.7 GB
20%
science
464,404
~1.5 GB
15%
Total
3,096,027
~12.7 GB
100%… See the full description on the dataset page: https://huggingface.co/datasets/Mkd-Yonas/keural-sft-mix.keural-conversation-ko
keural-conversation
한국어 일상 대화 AI 학습용 합성 데이터셋입니다.9개 주제, 1,350개 시나리오를 기반으로 191,093개의 멀티턴 대화 쌍을 생성했습니다.
데이터셋 개요
항목
내용
언어
한국어 (ko)
총 샘플 수
191,093개
시나리오 수
1,350개
대화 구조
2-turn (user → assistant → user → assistant)
생성 모델
google/gemma-4-26B-A4B-it
원본 생성 수
202,435개
수락률
99.97%
중복 제거율
5.60%
라이선스
CC BY 4.0
주제별 분포
주제
시나리오 수
샘플 수
자기소개 및 인사
150개
20,815
감정 표현과 공감
150개
21,411
날씨, 계절, 일상 근황
150개
20,855
취미, 관심사, 좋아하는 것
150개… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-conversation-ko.keural-rag-ko
keural-synthetic-SFT
한국어 RAG 특화 SFT(Supervised Fine-Tuning) 학습용 합성 데이터셋입니다.한국어 웹 문서 155만 개를 시드로 삼아 5가지 태스크 유형으로 1,402,611개의 질문-답변 쌍을 생성했습니다.
⚠️ RAG 특화 데이터셋 안내본 데이터셋의 답변은 시드 문서(컨텍스트)를 기반으로 생성되었습니다.답변 내에 "제시된 텍스트를 바탕으로", "문서에 따르면" 등의 표현이 포함될 수 있으며,RAG(Retrieval-Augmented Generation) 파인튜닝 또는 문서 기반 QA 모델 학습에 최적화되어 있습니다.컨텍스트 없이 일반 지식 Q&A 모델을 학습하려는 경우에는 별도의 general SFT 데이터셋을 권장합니다.
데이터셋 개요
항목
내용
언어
한국어 (ko)
총 샘플 수
1,402,611개
시드 문서 수
1,545,735개 (3,436개 고유… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-rag-ko.keural-nova-v1.2-sft
Keural Nova v1.2 — SFT dataset (PRIVATE)
The supervised fine-tuning mix used to train Keural Nova v1.2. Cleaned & balanced:
benchmark test-splits excluded, code AST-validated, identity de-contaminated,
AI-Hub / non-commercial rows removed. Each row carries accurate source_name and license fields.
Total 146,274 rows — train 144,812 / eval 1,462.
Format: {"messages":[{"role","content"},...], "source_name":..., "license":...}
Composition (by source)… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-v1.2-sft.keural-DPO
keural-DPO
한국어/영어 DPO(Direct Preference Optimization) 학습용 통합 데이터셋.
6개 소스 데이터셋을 ChatML 형식으로 정규화하여 통합한 것으로, 총 443,314 샘플을 포함합니다.
데이터셋 통계
Split
소스
필터 제거
업로드
언어
라이선스
ultrafeedback_binarized
61,054
424
60,630
EN
MIT
multifaceted_collection_dpo
65,139
1,045
64,094
EN/KO
Apache-2.0
hh_rlhf
160,608
739
159,869
EN
MIT
orca_dpo_pairs_ko
12,727
2
12,725
KO
—
aihub_71748
29,676
0
29,676
KO
AI Hub
aihub_71760
116,320
0
116,320
KO
AI Hub
합계
445,524
2,210
443,314… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-DPO.keural-cortex-8b-sft
Keural-Cortex-8B SFT dataset
The supervised fine-tuning set used to train Keural-Cortex-8B, a Korean-first
bilingual model with a 64K context window, tool calling, and hybrid
thinking/non-thinking modes.
1,568,649 rows · 1.87B estimated tokens · 1.80B real Qwen3 tokens · 73.7% Korean
Three files:
file
rows
what it is
train.jsonl
1,564,042
main set, all rows under 32,768 tokens
train_long64k.jsonl
4,607
the 32K–64K band, kept separate because it needs a different… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-cortex-8b-sft.keural-v2-tool-calling
Tool & Function Calling (Area 2) — Korean SFT Dataset Prep
상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님
출처
원본: glaiveai/glaive-function-calling-v2
커밋 해시: e7f4b6456019f5d8bcb991ef0dd67d8ff23221ac
라이선스: Apache-2.0 (원본 태그, README 본문 없어 대조 문구 없음)
생성 출처: 미확인 — GPT-4/Claude 등 프론티어 모델 사용 가능성 있음 (원본 데이터셋 카드에 명시 없음)
언어: 영어 (지침서 §1.2 정책에 따라 번역 없이 영어 그대로 사용)
처리 과정
원본 112,960건 다운로드
chat 필드 기준 완전 중복 23,790건(21%) 발견 및 제거 → 유니크 89,170건
유니크 풀에서 seed=42로 50… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-tool-calling.keural-v2-cot-reasoning
Reasoning / Chain-of-Thought (Area 4) — Korean SFT Dataset Prep
상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님
출처
원본: nvidia/OpenMathReasoning (cot split)
커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df
라이선스: CC-BY-4.0 (태그와 본문 일치, "License/Terms of Use: cc-by-4.0")
생성 모델: DeepSeek-R1(샘플 중 다수), QwQ-32B — 둘 다 오픈 웨이트 모델, 독점 모델 ToS 리스크 없음
언어: 영어 (지침서 §1.4 정책에 따라 번역 없이 영어 그대로 사용)
출처 구성 (problem_source)
문제(질문) 출처는 대부분 AoPS(Art of Problem Solving) 포럼… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning.keural-v2-self-verification
Self-Verification Dataset
Status: generation in progress. 4,477 / 50,000 target rows (~9%), growing. Being generated in parallel across multiple environments/models — see Generation models mix below. This card describes the file as of this snapshot; row count and model mix will change on re-upload.
File: 05_self_verification_generated.jsonl (one JSON object per line).
What this is
Synthetic "wrong draft → self-critique → corrected solution" traces, for training a… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-self-verification.keural-v2-orchestration
Agent Orchestration Dataset
Status: generation complete. 50,000 / 50,000 target rows. Not yet processed — see Pipeline stage before using this for training.
File: agent_orchestration_final_50000.jsonl (597 MB, one JSON object per line, 50,000 lines).
What this is
Synthetic multi-agent conversation traces for training a model to act as an orchestrator: decompose a user task, delegate subtasks to named sub-agents, then synthesize their independent responses into one… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-orchestration.keural-v2-fluency-v2
Keural-v2 Fluency (v2): A Curated Korean–English Conversational Corpus for LLM Fine-Tuning
Introduction |
Dataset Composition |
Methodology |
License |
Limitations
Status: Private staging — full §3 processing pipeline complete (dedup → PII removal → Korean-purity filter → length check → ratio measurement → train/val/test split → schema normalization → target-model encoding). Pending §4 quantitative/qualitative evaluation and second-party license audit before any… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency-v2.keural-QA-en
keural-synthetic-SFT-en-general
English general-purpose SFT dataset synthetically generated from English Wikipedia using a two-stage pipeline. Unlike RAG-specific datasets, questions and answers contain no references to source documents — the model is expected to answer from its own knowledge, with Wikipedia serving only as a grounding anchor during generation.
Dataset Summary
Item
Value
Total accepted
~1.6M
Total rejected
~8,700
Accept rate
~99.5%
Seed… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-QA-en.keural-DPO-raw
keural-DPO-raw
keural-DPO의 전처리 전 normalized 원본 데이터.
필터링 없이 원본 그대로 보존하여 재현성을 위해 제공합니다.
ChatML 변환 및 필터링이 적용된 학습용 데이터는 mkd-chanwoo/keural-DPO 를 사용하세요.
데이터셋 통계
Split
행 수
언어
라이선스
원본
ultrafeedback
61,054
EN
MIT
HuggingFaceH4/ultrafeedback_binarized
multifaceted_dpo
65,139
EN/KO
Apache-2.0
kaist-ai/Multifaceted-Collection-DPO
hh_rlhf
160,608
EN
MIT
Anthropic/hh-rlhf
orca_dpo_ko
12,727
KO
—
Ja-ck/Orca-DPO-Pairs-KO
aihub_71748
29,676
KO
AI Hub
AI Hub 71748… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-DPO-raw.keural-v2-cot-reasoning-v2
Reasoning / Chain-of-Thought (Area 4, v2) — Korean SFT Dataset Prep
상태: 비공개 스테이징(private) — §3 처리(1~8번, 최종 인코딩 포함) 전부 완료. 제2자 감사 전, 공개 배포 대상 아님.
이 v2는 §3 처리를 새로 검증하며 진행한 최종 버전입니다(2026-08-10). v1(원본 problem/generated_solution 필드 그대로)과 달리, DeepSeek-V4-Flash-0731 학습용 최종 텍스트(text 필드)로 인코딩까지 완료됐습니다.
출처
원본: nvidia/OpenMathReasoning (cot split)
커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df
라이선스: CC-BY-4.0 (태그와 본문 일치)
생성 모델: DeepSeek-R1(다수), QwQ-32B — 둘 다 오픈 웨이트 모델
언어:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning-v2.keural-datasets-sampleskeural-datasets-1kkeural-v2-tool-calling-v2
Tool & Function Calling (Area 2, v2) — Korean SFT Dataset Prep
상태: 비공개 스테이징(private) — §3 처리(1~6번, 스키마 정규화) 완료, §3-7(최종 텍스트 인코딩)만 보류. 제2자 감사 전, 공개 배포 대상 아님.
이 v2는 §3 처리를 새로 검증하며 발견한 오류를 수정한 버전입니다(2026-08-10). v1(원본 chat 텍스트 그대로)과 달리, 이 저장소엔 korean_sft_schema.md 통합 구조(source/license/lang/category/conversations:[{role,content,reasoning_content,tool_calls}])로 정규화된 데이터가 들어있습니다.
출처
원본: glaiveai/glaive-function-calling-v2
커밋 해시: e7f4b6456019f5d8bcb991ef0dd67d8ff23221ac… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-tool-calling-v2.keural-dpo-mergedkeural-nova-tooluse
Keural Nova — tool-calling SFT slice (PRIVATE)
Tool-calling data used for Keural Nova v1.2. 17,337 rows: single-turn, multi-turn
(call -> tool result -> final answer), negative (no-call), and long-context up to 32k tokens.
Rendered as Qwen XML tool calls via ms-swift's native agent schema
(tool_call/tool roles + per-row tools JSON string).
tooluse_short.jsonl — 14,337 rows (<= ~3.6k tokens)
tooluse_long32k.jsonl — 3,000 rows (8k–32k tokens)
Sources / licenses:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-tooluse.keural-nova-identity
Keural Nova — identity SFT data (PRIVATE)
MKD-original. 200 rows (ko+en) teaching the assistant it is Keural, developed by MKD,
including denials of Qwen / GPT / Gemini / Claude / Llama etc. Facts grounded in https://mkd.kr.
Format: {"messages":[{"role":"user"...},{"role":"assistant"...}]}. Mix with heavy replay when
training (identity-only over-fits). Released Apache-2.0.
