keural
Datasets
All datasets matching “keural”keural-datasets
Keural Pretraining Datasets (Stage 2)
Stage 2 final production corpus for training the Keural Korean LLM.
Quality-filtered, deduplicated, and domain-balanced across 4 domains.
Summary
Metric
Value
Total processed documents (post-filter)
757,710,609
Dedup removed (Stage 2)
93,919,634
Final documents
663,790,975
Total tokens
~522B
Domains
English, Korean, Code, Science
Source datasets
43
Format
Parquet (snappy compressed, sharded)
Upload… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-datasets.Keural-MoE-14B-stage1-Datasetkeural-SFT
Keural SFT Dataset
Bilingual (Korean/English) instruction-tuning dataset for the Keural LLM project.
Built from 14 curated sources and formatted in ChatML after multi-stage filtering.
Dataset Summary
Field
Value
Total samples
1,144,119
Total tokens
710,280,675 (~710M)
Average tokens/sample
621.0
Max sequence length
8,192 tokens
Language ratio
Korean 45.3% / English 54.7%
Format
ChatML
Number of shards
115 (10,000 samples/shard)
Tokenizer
Keural… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-SFT.keural-SFT-chatml-ko-v1
Keural SFT ChatML (Korean) v1
한국어 SFT(Supervised Fine-Tuning)용 통합 데이터셋입니다. 공개 한국어 instruction/대화 데이터셋 8종을 수집하여 정제 → 품질 필터링 → 안전성 필터링 → 중복 제거(exact + MinHash near-dup) → ChatML 포맷팅 → 토크나이즈 검증을 거친 결과물입니다.
총 샘플 수: 710,278
총 토큰 수: 약 1.9억 (190,724,939 tokens, keural tokenizer 기준)
포맷: ChatML (<|im_start|>role ... <|im_end|>)
최대 길이: 8,192 tokens (초과 시 truncate)
생성일: 2026-07-10
데이터 구조
각 샤드는 JSONL 형식이며, 레코드 스키마는 다음과 같습니다:
{"text":… See the full description on the dataset page: https://huggingface.co/datasets/mkd-jueon/keural-SFT-chatml-ko-v1.keural-SFT-chatml-ko-v3
Keural SFT ChatML — Korean
Korean supervised fine-tuning (SFT) dataset in fully-rendered ChatML format, produced by the
Keural SFT data pipeline (collect → structure → clean → quality-filter → safety-filter → dedup →
format → tokenize → package → audit). Intended for instruction-tuning of the Keural model family.
Language: Korean (ko)
Total samples: 2,018,250 (~499M tokens; see manifest.json for per-shard counts and sha256 checksums)
Format: one JSON object per line (JSONL)… See the full description on the dataset page: https://huggingface.co/datasets/Mkd-Yonas/keural-SFT-chatml-ko-v3.keural-SFT-rebuilt
Keural SFT — Rebuilt (indentation-fixed)
Rebuild of the keural-SFT mixed EN/KO instruction dataset, reprocessed from the original
public sources with a fixed cleaning stage.
⚠️ Why this rebuild exists: the original mkd-chanwoo/keural-SFT was produced with a
cleaner rule that collapsed consecutive spaces, which flattened all code-block indentation
to one space — code samples were syntactically broken (root cause of a HumanEval
regression). The corruption is not recoverable from… See the full description on the dataset page: https://huggingface.co/datasets/Mkd-Yonas/keural-SFT-rebuilt.
