CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkd-chanwoo /keural-datasets Keural Pretraining Datasets (Stage 2) Stage 2 final production corpus for training the Keural Korean LLM. Quality-filtered, deduplicated, and domain-balanced across 4 domains. Summary Metric Value Total processed documents (post-filter) 757,710,609 Dedup removed (Stage 2) 93,919,634 Final documents 663,790,975 Total tokens ~522B Domains English, Korean, Code, Science Source datasets 43 Format Parquet (snappy compressed, sharded) Upload… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-datasets.tabular100M<n<1B0 likes1.6k downloads4mo agoHugging Face02mkd-hossain /Keural-MoE-14B-stage1-Datasettext10M<n<100M0 likes863 downloads6mo agoHugging Face03mkd-jueon /keural-SFT-chatml-ko-v1 Keural SFT ChatML (Korean) v1 한국어 SFT(Supervised Fine-Tuning)용 통합 데이터셋입니다. 공개 한국어 instruction/대화 데이터셋 8종을 수집하여 정제 → 품질 필터링 → 안전성 필터링 → 중복 제거(exact + MinHash near-dup) → ChatML 포맷팅 → 토크나이즈 검증을 거친 결과물입니다. 총 샘플 수: 710,278 총 토큰 수: 약 1.9억 (190,724,939 tokens, keural tokenizer 기준) 포맷: ChatML (<|im_start|>role ... <|im_end|>) 최대 길이: 8,192 tokens (초과 시 truncate) 생성일: 2026-07-10 데이터 구조 각 샤드는 JSONL 형식이며, 레코드 스키마는 다음과 같습니다: {"text":… See the full description on the dataset page: https://huggingface.co/datasets/mkd-jueon/keural-SFT-chatml-ko-v1.texttext-generation1M<n<10M2 likes213 downloads3mo agoHugging Face04mkd-minju /keural-v2-dataset Keural v2 — MoE Fine-Tuning Dataset 상태: 비공개 (private) — 공개 배포 대상 아님 Keural MoE Pro v2 모델 파인튜닝을 위해 8개 카테고리(A~H)로 구성된 SFT 학습 데이터셋입니다. 자세한 수집·처리 기준은 mkd-minju/Keural-MoE-Pro-v2 GitHub 저장소의 docs/V2-DATASET-PREP.md 계획 문서를 따릅니다. 카테고리 구성 파일 카테고리 목표 건수 언어 주요 출처 라이선스 A_korean_conversation.jsonl 한국어 대화/지침 50,000 ko mkd-chanwoo/keural-conversation-chatml-ko, mkd-chanwoo/keural-rag-chatml-ko… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-dataset.texttext-generation100K<n<1M0 likes73 downloads24d agoHugging Face05mkd-chanwoo /keural-conversation-chatml-ko keural-conversation-chatml-ko mkd-chanwoo/keural-conversation-ko 데이터셋을 SFT 학습용 ChatML 포맷으로 전처리한 한국어 일상대화 데이터셋입니다. 9개 주제의 일상 캐주얼 대화를 Gemma-4-26B 모델이 생성한 synthetic 멀티턴 대화로 구성되어 있습니다. 데이터셋 개요 항목 값 총 샘플 수 133,339 총 토큰 수 약 19M 평균 토큰 / 샘플 146.0 중간값 토큰 145 최대 토큰 805 최소 토큰 95 언어 한국어 (100%) 포맷 ChatML 라이선스 CC BY 4.0 포맷 ChatML 형식의 멀티턴 대화입니다. <|im_start|>user {발화 1} <|im_end|> <|im_start|>assistant {발화 2} <|im_end|> <|im_start|>user {발화… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-conversation-chatml-ko.texttext-generation100K<n<1M0 likes66 downloads4mo agoHugging Face06Mkd-Yonas /keural-sft-mix Keural SFT Mix (sampled) A language/domain-balanced sample of mkd-chanwoo/keural-datasets, drawn for SFT experimentation. Each category is kept as its own config (mirroring the source layout), so the different category schemas never have to be merged. Composition Config Rows Size Target mix korean 1,238,411 ~6.3 GB 40% english 774,007 ~2.1 GB 25% code 619,205 ~2.7 GB 20% science 464,404 ~1.5 GB 15% Total 3,096,027 ~12.7 GB 100%… See the full description on the dataset page: https://huggingface.co/datasets/Mkd-Yonas/keural-sft-mix.tabular1M<n<10M0 likes64 downloads3mo agoHugging Face07mkd-chanwoo /keural-conversation-ko keural-conversation 한국어 일상 대화 AI 학습용 합성 데이터셋입니다.9개 주제, 1,350개 시나리오를 기반으로 191,093개의 멀티턴 대화 쌍을 생성했습니다. 데이터셋 개요 항목 내용 언어 한국어 (ko) 총 샘플 수 191,093개 시나리오 수 1,350개 대화 구조 2-turn (user → assistant → user → assistant) 생성 모델 google/gemma-4-26B-A4B-it 원본 생성 수 202,435개 수락률 99.97% 중복 제거율 5.60% 라이선스 CC BY 4.0 주제별 분포 주제 시나리오 수 샘플 수 자기소개 및 인사 150개 20,815 감정 표현과 공감 150개 21,411 날씨, 계절, 일상 근황 150개 20,855 취미, 관심사, 좋아하는 것 150개… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-conversation-ko.texttext-generation100K<n<1M0 likes62 downloads4mo agoHugging Face08mkd-chanwoo /keural-rag-ko keural-synthetic-SFT 한국어 RAG 특화 SFT(Supervised Fine-Tuning) 학습용 합성 데이터셋입니다.한국어 웹 문서 155만 개를 시드로 삼아 5가지 태스크 유형으로 1,402,611개의 질문-답변 쌍을 생성했습니다. ⚠️ RAG 특화 데이터셋 안내본 데이터셋의 답변은 시드 문서(컨텍스트)를 기반으로 생성되었습니다.답변 내에 "제시된 텍스트를 바탕으로", "문서에 따르면" 등의 표현이 포함될 수 있으며,RAG(Retrieval-Augmented Generation) 파인튜닝 또는 문서 기반 QA 모델 학습에 최적화되어 있습니다.컨텍스트 없이 일반 지식 Q&A 모델을 학습하려는 경우에는 별도의 general SFT 데이터셋을 권장합니다. 데이터셋 개요 항목 내용 언어 한국어 (ko) 총 샘플 수 1,402,611개 시드 문서 수 1,545,735개 (3,436개 고유… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-rag-ko.texttext-generation1M<n<10M0 likes53 downloads4mo agoHugging Face09mkd-hossain /keural-nova-v1.2-sft Keural Nova v1.2 — SFT dataset (PRIVATE) The supervised fine-tuning mix used to train Keural Nova v1.2. Cleaned & balanced: benchmark test-splits excluded, code AST-validated, identity de-contaminated, AI-Hub / non-commercial rows removed. Each row carries accurate source_name and license fields. Total 146,274 rows — train 144,812 / eval 1,462. Format: {"messages":[{"role","content"},...], "source_name":..., "license":...} Composition (by source)… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-v1.2-sft.texttext-generation100K<n<1M1 likes51 downloads2mo agoHugging Face10mkd-chanwoo /keural-DPO keural-DPO 한국어/영어 DPO(Direct Preference Optimization) 학습용 통합 데이터셋. 6개 소스 데이터셋을 ChatML 형식으로 정규화하여 통합한 것으로, 총 443,314 샘플을 포함합니다. 데이터셋 통계 Split 소스 필터 제거 업로드 언어 라이선스 ultrafeedback_binarized 61,054 424 60,630 EN MIT multifaceted_collection_dpo 65,139 1,045 64,094 EN/KO Apache-2.0 hh_rlhf 160,608 739 159,869 EN MIT orca_dpo_pairs_ko 12,727 2 12,725 KO — aihub_71748 29,676 0 29,676 KO AI Hub aihub_71760 116,320 0 116,320 KO AI Hub 합계 445,524 2,210 443,314… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-DPO.texttext-generation100K<n<1M0 likes48 downloads4mo agoHugging Face11mkd-hossain /keural-cortex-8b-sft Keural-Cortex-8B SFT dataset The supervised fine-tuning set used to train Keural-Cortex-8B, a Korean-first bilingual model with a 64K context window, tool calling, and hybrid thinking/non-thinking modes. 1,568,649 rows · 1.87B estimated tokens · 1.80B real Qwen3 tokens · 73.7% Korean Three files: file rows what it is train.jsonl 1,564,042 main set, all rows under 32,768 tokens train_long64k.jsonl 4,607 the 32K–64K band, kept separate because it needs a different… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-cortex-8b-sft.texttext-generation10K<n<100K0 likes46 downloads8d agoHugging Face12mkd-minju /keural-v2-tool-calling Tool & Function Calling (Area 2) — Korean SFT Dataset Prep 상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님 출처 원본: glaiveai/glaive-function-calling-v2 커밋 해시: e7f4b6456019f5d8bcb991ef0dd67d8ff23221ac 라이선스: Apache-2.0 (원본 태그, README 본문 없어 대조 문구 없음) 생성 출처: 미확인 — GPT-4/Claude 등 프론티어 모델 사용 가능성 있음 (원본 데이터셋 카드에 명시 없음) 언어: 영어 (지침서 §1.2 정책에 따라 번역 없이 영어 그대로 사용) 처리 과정 원본 112,960건 다운로드 chat 필드 기준 완전 중복 23,790건(21%) 발견 및 제거 → 유니크 89,170건 유니크 풀에서 seed=42로 50… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-tool-calling.texttext-generation10K<n<100K0 likes32 downloads24d agoHugging Face13mkd-minju /keural-v2-cot-reasoning Reasoning / Chain-of-Thought (Area 4) — Korean SFT Dataset Prep 상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님 출처 원본: nvidia/OpenMathReasoning (cot split) 커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df 라이선스: CC-BY-4.0 (태그와 본문 일치, "License/Terms of Use: cc-by-4.0") 생성 모델: DeepSeek-R1(샘플 중 다수), QwQ-32B — 둘 다 오픈 웨이트 모델, 독점 모델 ToS 리스크 없음 언어: 영어 (지침서 §1.4 정책에 따라 번역 없이 영어 그대로 사용) 출처 구성 (problem_source) 문제(질문) 출처는 대부분 AoPS(Art of Problem Solving) 포럼… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning.texttext-generation10K<n<100K0 likes31 downloads24d agoHugging Face14mkd-minju /keural-v2-self-verification Self-Verification Dataset Status: generation in progress. 4,477 / 50,000 target rows (~9%), growing. Being generated in parallel across multiple environments/models — see Generation models mix below. This card describes the file as of this snapshot; row count and model mix will change on re-upload. File: 05_self_verification_generated.jsonl (one JSON object per line). What this is Synthetic "wrong draft → self-critique → corrected solution" traces, for training a… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-self-verification.texttext-generation1K<n<10K0 likes29 downloads23d agoHugging Face15mkd-minju /keural-v2-orchestration Agent Orchestration Dataset Status: generation complete. 50,000 / 50,000 target rows. Not yet processed — see Pipeline stage before using this for training. File: agent_orchestration_final_50000.jsonl (597 MB, one JSON object per line, 50,000 lines). What this is Synthetic multi-agent conversation traces for training a model to act as an orchestrator: decompose a user task, delegate subtasks to named sub-agents, then synthesize their independent responses into one… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-orchestration.texttext-generation10K<n<100K0 likes27 downloads24d agoHugging Face16mkd-minju /keural-v2-fluency-v2 Keural-v2 Fluency (v2): A Curated Korean–English Conversational Corpus for LLM Fine-Tuning Introduction | Dataset Composition | Methodology | License | Limitations Status: Private staging — full §3 processing pipeline complete (dedup → PII removal → Korean-purity filter → length check → ratio measurement → train/val/test split → schema normalization → target-model encoding). Pending §4 quantitative/qualitative evaluation and second-party license audit before any… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency-v2.texttext-generation100K<n<1M0 likes21 downloads1mo agoHugging Face17mkd-chanwoo /keural-QA-en keural-synthetic-SFT-en-general English general-purpose SFT dataset synthetically generated from English Wikipedia using a two-stage pipeline. Unlike RAG-specific datasets, questions and answers contain no references to source documents — the model is expected to answer from its own knowledge, with Wikipedia serving only as a grounding anchor during generation. Dataset Summary Item Value Total accepted ~1.6M Total rejected ~8,700 Accept rate ~99.5% Seed… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-QA-en.textquestion-answering1M<n<10M0 likes20 downloads4mo agoHugging Face18mkd-chanwoo /keural-DPO-raw keural-DPO-raw keural-DPO의 전처리 전 normalized 원본 데이터. 필터링 없이 원본 그대로 보존하여 재현성을 위해 제공합니다. ChatML 변환 및 필터링이 적용된 학습용 데이터는 mkd-chanwoo/keural-DPO 를 사용하세요. 데이터셋 통계 Split 행 수 언어 라이선스 원본 ultrafeedback 61,054 EN MIT HuggingFaceH4/ultrafeedback_binarized multifaceted_dpo 65,139 EN/KO Apache-2.0 kaist-ai/Multifaceted-Collection-DPO hh_rlhf 160,608 EN MIT Anthropic/hh-rlhf orca_dpo_ko 12,727 KO — Ja-ck/Orca-DPO-Pairs-KO aihub_71748 29,676 KO AI Hub AI Hub 71748… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-DPO-raw.texttext-generation100K<n<1M0 likes17 downloads4mo agoHugging Face19mkd-minju /keural-v2-cot-reasoning-v2 Reasoning / Chain-of-Thought (Area 4, v2) — Korean SFT Dataset Prep 상태: 비공개 스테이징(private) — §3 처리(1~8번, 최종 인코딩 포함) 전부 완료. 제2자 감사 전, 공개 배포 대상 아님. 이 v2는 §3 처리를 새로 검증하며 진행한 최종 버전입니다(2026-08-10). v1(원본 problem/generated_solution 필드 그대로)과 달리, DeepSeek-V4-Flash-0731 학습용 최종 텍스트(text 필드)로 인코딩까지 완료됐습니다. 출처 원본: nvidia/OpenMathReasoning (cot split) 커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df 라이선스: CC-BY-4.0 (태그와 본문 일치) 생성 모델: DeepSeek-R1(다수), QwQ-32B — 둘 다 오픈 웨이트 모델 언어:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning-v2.texttext-generation10K<n<100K0 likes16 downloads1mo agoHugging Face20mkd-chanwoo /keural-datasets-samplestabular100K<n<1M0 likes15 downloads4mo agoHugging Face21mkd-jueon /keural-datasets-1ktext1K<n<10K0 likes15 downloads3mo agoHugging Face22mkd-minju /keural-v2-tool-calling-v2 Tool & Function Calling (Area 2, v2) — Korean SFT Dataset Prep 상태: 비공개 스테이징(private) — §3 처리(1~6번, 스키마 정규화) 완료, §3-7(최종 텍스트 인코딩)만 보류. 제2자 감사 전, 공개 배포 대상 아님. 이 v2는 §3 처리를 새로 검증하며 발견한 오류를 수정한 버전입니다(2026-08-10). v1(원본 chat 텍스트 그대로)과 달리, 이 저장소엔 korean_sft_schema.md 통합 구조(source/license/lang/category/conversations:[{role,content,reasoning_content,tool_calls}])로 정규화된 데이터가 들어있습니다. 출처 원본: glaiveai/glaive-function-calling-v2 커밋 해시: e7f4b6456019f5d8bcb991ef0dd67d8ff23221ac… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-tool-calling-v2.texttext-generation10K<n<100K0 likes15 downloads1mo agoHugging Face23mkd-hossain /keural-dpo-mergedtext100K<n<1M0 likes9 downloads4mo agoHugging Face24mkd-hossain /keural-nova-tooluse Keural Nova — tool-calling SFT slice (PRIVATE) Tool-calling data used for Keural Nova v1.2. 17,337 rows: single-turn, multi-turn (call -> tool result -> final answer), negative (no-call), and long-context up to 32k tokens. Rendered as Qwen XML tool calls via ms-swift's native agent schema (tool_call/tool roles + per-row tools JSON string). tooluse_short.jsonl — 14,337 rows (<= ~3.6k tokens) tooluse_long32k.jsonl — 3,000 rows (8k–32k tokens) Sources / licenses:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-tooluse.texttext-generation10K<n<100K0 likes5 downloads2mo agoHugging Face25mkd-hossain /keural-nova-identity Keural Nova — identity SFT data (PRIVATE) MKD-original. 200 rows (ko+en) teaching the assistant it is Keural, developed by MKD, including denials of Qwen / GPT / Gemini / Claude / Llama etc. Facts grounded in https://mkd.kr. Format: {"messages":[{"role":"user"...},{"role":"assistant"...}]}. Mix with heavy replay when training (identity-only over-fits). Released Apache-2.0. texttext-generationn<1K0 likes2 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.