CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkd-jueon /keural-SFT-chatml-ko-v1 Keural SFT ChatML (Korean) v1 한국어 SFT(Supervised Fine-Tuning)용 통합 데이터셋입니다. 공개 한국어 instruction/대화 데이터셋 8종을 수집하여 정제 → 품질 필터링 → 안전성 필터링 → 중복 제거(exact + MinHash near-dup) → ChatML 포맷팅 → 토크나이즈 검증을 거친 결과물입니다. 총 샘플 수: 710,278 총 토큰 수: 약 1.9억 (190,724,939 tokens, keural tokenizer 기준) 포맷: ChatML (<|im_start|>role ... <|im_end|>) 최대 길이: 8,192 tokens (초과 시 truncate) 생성일: 2026-07-10 데이터 구조 각 샤드는 JSONL 형식이며, 레코드 스키마는 다음과 같습니다: {"text":… See the full description on the dataset page: https://huggingface.co/datasets/mkd-jueon/keural-SFT-chatml-ko-v1.texttext-generation1M<n<10M2 likes213 downloads3mo agoHugging Face02mkd-minju /keural-v2-dataset Keural v2 — MoE Fine-Tuning Dataset 상태: 비공개 (private) — 공개 배포 대상 아님 Keural MoE Pro v2 모델 파인튜닝을 위해 8개 카테고리(A~H)로 구성된 SFT 학습 데이터셋입니다. 자세한 수집·처리 기준은 mkd-minju/Keural-MoE-Pro-v2 GitHub 저장소의 docs/V2-DATASET-PREP.md 계획 문서를 따릅니다. 카테고리 구성 파일 카테고리 목표 건수 언어 주요 출처 라이선스 A_korean_conversation.jsonl 한국어 대화/지침 50,000 ko mkd-chanwoo/keural-conversation-chatml-ko, mkd-chanwoo/keural-rag-chatml-ko… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-dataset.texttext-generation100K<n<1M0 likes73 downloads24d agoHugging Face03mkd-chanwoo /keural-conversation-chatml-ko keural-conversation-chatml-ko mkd-chanwoo/keural-conversation-ko 데이터셋을 SFT 학습용 ChatML 포맷으로 전처리한 한국어 일상대화 데이터셋입니다. 9개 주제의 일상 캐주얼 대화를 Gemma-4-26B 모델이 생성한 synthetic 멀티턴 대화로 구성되어 있습니다. 데이터셋 개요 항목 값 총 샘플 수 133,339 총 토큰 수 약 19M 평균 토큰 / 샘플 146.0 중간값 토큰 145 최대 토큰 805 최소 토큰 95 언어 한국어 (100%) 포맷 ChatML 라이선스 CC BY 4.0 포맷 ChatML 형식의 멀티턴 대화입니다. <|im_start|>user {발화 1} <|im_end|> <|im_start|>assistant {발화 2} <|im_end|> <|im_start|>user {발화… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-conversation-chatml-ko.texttext-generation100K<n<1M0 likes66 downloads4mo agoHugging Face04mkd-hossain /keural-nova-v1.2-sft Keural Nova v1.2 — SFT dataset (PRIVATE) The supervised fine-tuning mix used to train Keural Nova v1.2. Cleaned & balanced: benchmark test-splits excluded, code AST-validated, identity de-contaminated, AI-Hub / non-commercial rows removed. Each row carries accurate source_name and license fields. Total 146,274 rows — train 144,812 / eval 1,462. Format: {"messages":[{"role","content"},...], "source_name":..., "license":...} Composition (by source)… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-v1.2-sft.texttext-generation100K<n<1M1 likes51 downloads2mo agoHugging Face05mkd-chanwoo /keural-DPO keural-DPO 한국어/영어 DPO(Direct Preference Optimization) 학습용 통합 데이터셋. 6개 소스 데이터셋을 ChatML 형식으로 정규화하여 통합한 것으로, 총 443,314 샘플을 포함합니다. 데이터셋 통계 Split 소스 필터 제거 업로드 언어 라이선스 ultrafeedback_binarized 61,054 424 60,630 EN MIT multifaceted_collection_dpo 65,139 1,045 64,094 EN/KO Apache-2.0 hh_rlhf 160,608 739 159,869 EN MIT orca_dpo_pairs_ko 12,727 2 12,725 KO — aihub_71748 29,676 0 29,676 KO AI Hub aihub_71760 116,320 0 116,320 KO AI Hub 합계 445,524 2,210 443,314… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-DPO.texttext-generation100K<n<1M0 likes48 downloads4mo agoHugging Face06mkd-hossain /keural-cortex-8b-sft Keural-Cortex-8B SFT dataset The supervised fine-tuning set used to train Keural-Cortex-8B, a Korean-first bilingual model with a 64K context window, tool calling, and hybrid thinking/non-thinking modes. 1,568,649 rows · 1.87B estimated tokens · 1.80B real Qwen3 tokens · 73.7% Korean Three files: file rows what it is train.jsonl 1,564,042 main set, all rows under 32,768 tokens train_long64k.jsonl 4,607 the 32K–64K band, kept separate because it needs a different… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-cortex-8b-sft.texttext-generation10K<n<100K0 likes46 downloads8d agoHugging Face07mkd-minju /keural-v2-tool-calling Tool & Function Calling (Area 2) — Korean SFT Dataset Prep 상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님 출처 원본: glaiveai/glaive-function-calling-v2 커밋 해시: e7f4b6456019f5d8bcb991ef0dd67d8ff23221ac 라이선스: Apache-2.0 (원본 태그, README 본문 없어 대조 문구 없음) 생성 출처: 미확인 — GPT-4/Claude 등 프론티어 모델 사용 가능성 있음 (원본 데이터셋 카드에 명시 없음) 언어: 영어 (지침서 §1.2 정책에 따라 번역 없이 영어 그대로 사용) 처리 과정 원본 112,960건 다운로드 chat 필드 기준 완전 중복 23,790건(21%) 발견 및 제거 → 유니크 89,170건 유니크 풀에서 seed=42로 50… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-tool-calling.texttext-generation10K<n<100K0 likes32 downloads24d agoHugging Face08mkd-minju /keural-v2-cot-reasoning Reasoning / Chain-of-Thought (Area 4) — Korean SFT Dataset Prep 상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님 출처 원본: nvidia/OpenMathReasoning (cot split) 커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df 라이선스: CC-BY-4.0 (태그와 본문 일치, "License/Terms of Use: cc-by-4.0") 생성 모델: DeepSeek-R1(샘플 중 다수), QwQ-32B — 둘 다 오픈 웨이트 모델, 독점 모델 ToS 리스크 없음 언어: 영어 (지침서 §1.4 정책에 따라 번역 없이 영어 그대로 사용) 출처 구성 (problem_source) 문제(질문) 출처는 대부분 AoPS(Art of Problem Solving) 포럼… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning.texttext-generation10K<n<100K0 likes31 downloads24d agoHugging Face09mkd-minju /keural-v2-self-verification Self-Verification Dataset Status: generation in progress. 4,477 / 50,000 target rows (~9%), growing. Being generated in parallel across multiple environments/models — see Generation models mix below. This card describes the file as of this snapshot; row count and model mix will change on re-upload. File: 05_self_verification_generated.jsonl (one JSON object per line). What this is Synthetic "wrong draft → self-critique → corrected solution" traces, for training a… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-self-verification.texttext-generation1K<n<10K0 likes29 downloads23d agoHugging Face10mkd-minju /keural-v2-orchestration Agent Orchestration Dataset Status: generation complete. 50,000 / 50,000 target rows. Not yet processed — see Pipeline stage before using this for training. File: agent_orchestration_final_50000.jsonl (597 MB, one JSON object per line, 50,000 lines). What this is Synthetic multi-agent conversation traces for training a model to act as an orchestrator: decompose a user task, delegate subtasks to named sub-agents, then synthesize their independent responses into one… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-orchestration.texttext-generation10K<n<100K0 likes27 downloads24d agoHugging Face11mkd-chanwoo /keural-DPO-raw keural-DPO-raw keural-DPO의 전처리 전 normalized 원본 데이터. 필터링 없이 원본 그대로 보존하여 재현성을 위해 제공합니다. ChatML 변환 및 필터링이 적용된 학습용 데이터는 mkd-chanwoo/keural-DPO 를 사용하세요. 데이터셋 통계 Split 행 수 언어 라이선스 원본 ultrafeedback 61,054 EN MIT HuggingFaceH4/ultrafeedback_binarized multifaceted_dpo 65,139 EN/KO Apache-2.0 kaist-ai/Multifaceted-Collection-DPO hh_rlhf 160,608 EN MIT Anthropic/hh-rlhf orca_dpo_ko 12,727 KO — Ja-ck/Orca-DPO-Pairs-KO aihub_71748 29,676 KO AI Hub AI Hub 71748… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-DPO-raw.texttext-generation100K<n<1M0 likes17 downloads4mo agoHugging Face12mkd-minju /keural-v2-cot-reasoning-v2 Reasoning / Chain-of-Thought (Area 4, v2) — Korean SFT Dataset Prep 상태: 비공개 스테이징(private) — §3 처리(1~8번, 최종 인코딩 포함) 전부 완료. 제2자 감사 전, 공개 배포 대상 아님. 이 v2는 §3 처리를 새로 검증하며 진행한 최종 버전입니다(2026-08-10). v1(원본 problem/generated_solution 필드 그대로)과 달리, DeepSeek-V4-Flash-0731 학습용 최종 텍스트(text 필드)로 인코딩까지 완료됐습니다. 출처 원본: nvidia/OpenMathReasoning (cot split) 커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df 라이선스: CC-BY-4.0 (태그와 본문 일치) 생성 모델: DeepSeek-R1(다수), QwQ-32B — 둘 다 오픈 웨이트 모델 언어:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning-v2.texttext-generation10K<n<100K0 likes16 downloads1mo agoHugging Face13mkd-chanwoo /keural-datasets-samplestabular100K<n<1M0 likes15 downloads4mo agoHugging Face14mkd-minju /keural-v2-tool-calling-v2 Tool & Function Calling (Area 2, v2) — Korean SFT Dataset Prep 상태: 비공개 스테이징(private) — §3 처리(1~6번, 스키마 정규화) 완료, §3-7(최종 텍스트 인코딩)만 보류. 제2자 감사 전, 공개 배포 대상 아님. 이 v2는 §3 처리를 새로 검증하며 발견한 오류를 수정한 버전입니다(2026-08-10). v1(원본 chat 텍스트 그대로)과 달리, 이 저장소엔 korean_sft_schema.md 통합 구조(source/license/lang/category/conversations:[{role,content,reasoning_content,tool_calls}])로 정규화된 데이터가 들어있습니다. 출처 원본: glaiveai/glaive-function-calling-v2 커밋 해시: e7f4b6456019f5d8bcb991ef0dd67d8ff23221ac… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-tool-calling-v2.texttext-generation10K<n<100K0 likes15 downloads1mo agoHugging Face15mkd-hossain /keural-nova-tooluse Keural Nova — tool-calling SFT slice (PRIVATE) Tool-calling data used for Keural Nova v1.2. 17,337 rows: single-turn, multi-turn (call -> tool result -> final answer), negative (no-call), and long-context up to 32k tokens. Rendered as Qwen XML tool calls via ms-swift's native agent schema (tool_call/tool roles + per-row tools JSON string). tooluse_short.jsonl — 14,337 rows (<= ~3.6k tokens) tooluse_long32k.jsonl — 3,000 rows (8k–32k tokens) Sources / licenses:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-nova-tooluse.texttext-generation10K<n<100K0 likes5 downloads2mo agoHugging Face16mkd-hossain /keural-nova-identity Keural Nova — identity SFT data (PRIVATE) MKD-original. 200 rows (ko+en) teaching the assistant it is Keural, developed by MKD, including denials of Qwen / GPT / Gemini / Claude / Llama etc. Facts grounded in https://mkd.kr. Format: {"messages":[{"role":"user"...},{"role":"assistant"...}]}. Mix with heavy replay when training (identity-only over-fits). Released Apache-2.0. texttext-generationn<1K0 likes2 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.