datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
keural-datasets
Keural Pretraining Datasets (Stage 2)
Stage 2 final production corpus for training the Keural Korean LLM.
Quality-filtered, deduplicated, and domain-balanced across 4 domains.
Summary
Metric
Value
Total processed documents (post-filter)
757,710,609
Dedup removed (Stage 2)
93,919,634
Final documents
663,790,975
Total tokens
~522B
Domains
English, Korean, Code, Science
Source datasets
43
Format
Parquet (snappy compressed, sharded)
Upload… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-datasets.keural-sft-mix
Keural SFT Mix (sampled)
A language/domain-balanced sample of mkd-chanwoo/keural-datasets,
drawn for SFT experimentation. Each category is kept as its own config (mirroring the source layout),
so the different category schemas never have to be merged.
Composition
Config
Rows
Size
Target mix
korean
1,238,411
~6.3 GB
40%
english
774,007
~2.1 GB
25%
code
619,205
~2.7 GB
20%
science
464,404
~1.5 GB
15%
Total
3,096,027
~12.7 GB
100%… See the full description on the dataset page: https://huggingface.co/datasets/Mkd-Yonas/keural-sft-mix.keural-datasets-samples
