CoolFace
14 results

keural

mkd-chanwoo /keural-datasets Keural Pretraining Datasets (Stage 2) Stage 2 final production corpus for training the Keural Korean LLM. Quality-filtered, deduplicated, and domain-balanced across 4 domains. Summary Metric Value Total processed documents (post-filter) 757,710,609 Dedup removed (Stage 2) 93,919,634 Final documents 663,790,975 Total tokens ~522B Domains English, Korean, Code, Science Source datasets 43 Format Parquet (snappy compressed, sharded) Upload… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-datasets.tabular100M<n<1B0 likes1.7k downloads4mo agoHugging Facemkd-hossain /Keural-MoE-14B-stage1-Datasettext10M<n<100M0 likes860 downloads6mo agoHugging Facemkd-chanwoo /keural-SFT Keural SFT Dataset Bilingual (Korean/English) instruction-tuning dataset for the Keural LLM project. Built from 14 curated sources and formatted in ChatML after multi-stage filtering. Dataset Summary Field Value Total samples 1,144,119 Total tokens 710,280,675 (~710M) Average tokens/sample 621.0 Max sequence length 8,192 tokens Language ratio Korean 45.3% / English 54.7% Format ChatML Number of shards 115 (10,000 samples/shard) Tokenizer Keural… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-SFT.1M<n<10M0 likes204 downloads5mo agoHugging Facemkd-jueon /keural-SFT-chatml-ko-v1 Keural SFT ChatML (Korean) v1 한국어 SFT(Supervised Fine-Tuning)용 통합 데이터셋입니다. 공개 한국어 instruction/대화 데이터셋 8종을 수집하여 정제 → 품질 필터링 → 안전성 필터링 → 중복 제거(exact + MinHash near-dup) → ChatML 포맷팅 → 토크나이즈 검증을 거친 결과물입니다. 총 샘플 수: 710,278 총 토큰 수: 약 1.9억 (190,724,939 tokens, keural tokenizer 기준) 포맷: ChatML (<|im_start|>role ... <|im_end|>) 최대 길이: 8,192 tokens (초과 시 truncate) 생성일: 2026-07-10 데이터 구조 각 샤드는 JSONL 형식이며, 레코드 스키마는 다음과 같습니다: {"text":… See the full description on the dataset page: https://huggingface.co/datasets/mkd-jueon/keural-SFT-chatml-ko-v1.texttext-generation1M<n<10M2 likes202 downloads3mo agoHugging FaceMkd-Yonas /keural-SFT-chatml-ko-v3 Keural SFT ChatML — Korean Korean supervised fine-tuning (SFT) dataset in fully-rendered ChatML format, produced by the Keural SFT data pipeline (collect → structure → clean → quality-filter → safety-filter → dedup → format → tokenize → package → audit). Intended for instruction-tuning of the Keural model family. Language: Korean (ko) Total samples: 2,018,250 (~499M tokens; see manifest.json for per-shard counts and sha256 checksums) Format: one JSON object per line (JSONL)… See the full description on the dataset page: https://huggingface.co/datasets/Mkd-Yonas/keural-SFT-chatml-ko-v3.text-generation1M<n<10M0 likes161 downloads2mo agoHugging FaceMkd-Yonas /keural-SFT-rebuilt Keural SFT — Rebuilt (indentation-fixed) Rebuild of the keural-SFT mixed EN/KO instruction dataset, reprocessed from the original public sources with a fixed cleaning stage. ⚠️ Why this rebuild exists: the original mkd-chanwoo/keural-SFT was produced with a cleaner rule that collapsed consecutive spaces, which flattened all code-block indentation to one space — code samples were syntactically broken (root cause of a HumanEval regression). The corruption is not recoverable from… See the full description on the dataset page: https://huggingface.co/datasets/Mkd-Yonas/keural-SFT-rebuilt.text-generation1M<n<10M0 likes128 downloads2mo agoHugging Face