CoolFace
Datasetpublic

mkd-chanwoo/normalized-datasets-for-koreanLLM

Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.

sourceHugging Faceupdated 4mo agoView on Hugging Face
1likes2.9kdownloads
settings

This repository belongs to mkd-chanwoo on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenormalized-datasets-for-koreanLLM
visibilitypublic
licencenot set
gatedno
ownermkd-chanwoo
Account settings
mkd-chanwoo/normalized-datasets-for-koreanLLM · CoolFace