CoolFace
Datasetpublic

mkd-chanwoo/normalized-datasets-for-koreanLLM

Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.

sourceHugging Faceupdated 4mo agoView on Hugging Face
1likes3.5kdownloads
Dataset Card

Normalized Datasets for Korean LLM (Stage 0.5)

This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.

Quick Stats

MetricValue
Total Documents~944,545,628
Source Datasets42
FormatJSONL (one JSON object per line)
LanguagesEnglish, Korean
DomainsEnglish, Korean, Code, Science
Pipeline StageStage 0.5 (after download, before filtering)
Last Updated2026-06-02

Total Documents by Domain

DomainDocuments% of Total
English360,821,20638.2%
Korean254,422,15326.9%
Code136,706,18314.5%
Science192,596,08620.4%
Total944,545,628100%

What is Normalization?

This dataset standardizes 42 source datasets with different formats into a unified JSONL schema. Each source uses different field names (e.g., TEXT, plain_text, content), so normalization:

  1. 1.Reads each source in its original format
  2. 2.Extracts text from its specific field
  3. 3.Writes to unified schema with consistent field names
  4. 4.Adds metadata (domain, language, doc_id, license)

Does NOT: modify text, filter, deduplicate, re-encode, or translate

Unified Document Schema

json
{
  "doc_id":             "gutenberg_000000042",
  "source_name":        "gutenberg",
  "domain":             "english",
  "language":           "en",
  "text":               "The full original text of the document...",
  "url":                "https://source-url.com/page",
  "license":            "Public Domain",
  "source_file":        "data/raw/gutenberg/train-00000-of-00001.parquet",
  "source_index":       42,
  "timestamp":          "2026-03-15T08:22:11Z",
  "processing_version": "v2"
}

Source Datasets & Field Mappings

DatasetDomainHuggingFace RepoText Field
gutenbergenglishsedthh/gutenberg_englishTEXT
ccnewsenglishstanford-oval/ccnewsplain_text
falcon-refinedwebenglishtiiuae/falcon-refinedwebcontent
finewebenglishHuggingFaceFW/finewebtext
fineweb_eduenglishHuggingFaceFW/fineweb-edutext
hpltenglish1/2/3englishHPLT/HPLT2.0_cleanedtext
openwebtextenglishSkylion007/openwebtexttext
wikipediaenglishwikimedia/wikipediatext
namuwikikoreanheegyu/namuwiki-extractedtext
wikipedia_kokoreanlcw99/wikipedia-korean-20240501text
oscarkoonlykoreanlcw99/oscar-ko-onlytext
korean_webtextkoreanHAERAE-HUB/KOREAN-WEBTEXTtext
hplt_koreankoreanHPLT/HPLT2.0_cleanedtext
culturax_kokoreanuonlp/CulturaXtext
wanjuan_koreankoreanopendatalab/WanJuan-Koreancontent
fineweb2_koreankoreanHuggingFaceFW/fineweb-2text
c4_koreankoreanallenai/c4text
cc100documentskoreankoreansingletongue/cc100-documentstext
cc_koreankoreanchaannwooff/CC-koreantext
korean_webkoreanchaannwooff/korean-webtext
korean_lawkoreanchaannwooff/Korean_lawtext
dartdockoreanchaannwooff/Dartdoctext
aihub_modukoreanAIHub (aihubshell)JSONPath extraction
aihub_bookskoreanAIHub (aihubshell)JSONPath extraction
aihubonlinecolloquialkoreanAIHub (aihubshell)JSONPath extraction
aihubkoreancorpus_literaturekoreanAIHub (aihubshell)JSONPath extraction
github-top-codecoderonantakizawa/github-top-codecontent
starcoderdatacodebigcode/starcoderdatacontent
codeparrot_cleancodecodeparrot/codeparrot-cleancontent
thestackccodebigcode/the-stackcontent
thestackjavacodebigcode/the-stackcontent
thestackpythoncodebigcode/the-stackcontent
arxivscienceKiteFishAI/arxiv-tex-corpus-fulltext
open-web-mathscienceopen-web-math/open-web-mathtext
peS2oscienceallenai/peS2otext
algebraic_stacksciencetypeof/algebraic-stacktext
fiwebmath_4plusscienceHuggingFaceTB/finemathtext
openmedtextscienceywchoi/OpenMedTexttext
pubmed_abstractssciencecasinca/PUBMEDtitleabstracts2019baselinetext
s2orcsciencesentence-transformers/s2orcabstract

Normalization Statistics (Seen vs Written)

Datasets processed by the Stage 0.5 normalizer (tracked in checkpoint). Datasets not listed below were normalized via legacy path.

DatasetSeenWrittenRate
algebraic_stack3,440,6943,440,226100.0%
c4_korean15,618,71815,618,718100.0%
cc100documentskorean35,678,35835,678,358100.0%
cc_korean3,570,6753,570,675100.0%
culturax_ko20,557,31020,557,309100.0%
dartdoc256,548256,548100.0%
fineweb81,595,32481,595,324100.0%
fineweb2_korean60,904,42960,904,429100.0%
fineweb_edu57,121,16757,121,167100.0%
fiwebmath_4plus6,699,4936,699,493100.0%
hpltenglish137,128,24437,128,244100.0%
hpltenglish216,085,77916,085,779100.0%
hpltenglish322,951,51522,951,515100.0%
hplt_korean38,866,83538,866,835100.0%
korean_law2,399,0062,399,006100.0%
korean_web1,545,7351,545,735100.0%
korean_webtext1,284,8791,284,878100.0%
openmedtext127,736127,736100.0%
oscarkoonly3,675,4213,675,420100.0%
pubmed_abstracts15,518,00915,517,555100.0%
s2orc89,455,93989,455,939100.0%
thestackc21,383,83221,383,810100.0%
thestackjava42,429,21142,429,204100.0%
thestackpython24,214,27024,214,204100.0%
wanjuan_korean68,894,93868,894,937100.0%
wikipedia6,407,8146,407,814100.0%
aihubkoreancorpus_literature3,276,0579080.0% ⚠️
⚠️ aihub_korean_corpus_literature: raw JSON 구조 미스매치로 정규화율 매우 낮음. 추후 extraction path 수정 필요.

Key Features

  • Resumable Processing: Checkpoint system allows resuming interrupted normalization
  • High Normalization Rate: Most datasets achieve >99% normalization (minimal parsing failures)
  • Mixed Licenses: Includes public domain, CC, MIT, Apache, and restricted licenses
  • Same Tokenizer: Uses Keural SentencePiece tokenizer (mkd-ai/keural-tokenizer) for token counting

Pipeline Context

Stage 0: Raw Download (~1.5 TB)
    ↓
Stage 0.5: Normalization → ~944M docs (YOU ARE HERE)
    ↓
Stage 1: Filtering → ~704M docs
    ↓
Stage 2: Dedup + Shard → ~512B tokens

License Considerations

Commercial use varies by source:

  • OK: FineWeb, Wikipedia, OpenWebText, C4, arXiv, PubMed (with conditions)
  • ⚠️ Caution: CCNews, Falcon RefinedWeb, some medical sources
  • No Commercial: NamuWiki (CC BY-NC-SA), AIHub datasets (research only)
  • 🔒 Restricted: StarCoderData, The Stack (BigCode OpenRAIL-M)

Review individual source licenses before commercial use.