datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
keural-rag-ko
keural-synthetic-SFT
한국어 RAG 특화 SFT(Supervised Fine-Tuning) 학습용 합성 데이터셋입니다.한국어 웹 문서 155만 개를 시드로 삼아 5가지 태스크 유형으로 1,402,611개의 질문-답변 쌍을 생성했습니다.
⚠️ RAG 특화 데이터셋 안내본 데이터셋의 답변은 시드 문서(컨텍스트)를 기반으로 생성되었습니다.답변 내에 "제시된 텍스트를 바탕으로", "문서에 따르면" 등의 표현이 포함될 수 있으며,RAG(Retrieval-Augmented Generation) 파인튜닝 또는 문서 기반 QA 모델 학습에 최적화되어 있습니다.컨텍스트 없이 일반 지식 Q&A 모델을 학습하려는 경우에는 별도의 general SFT 데이터셋을 권장합니다.
데이터셋 개요
항목
내용
언어
한국어 (ko)
총 샘플 수
1,402,611개
시드 문서 수
1,545,735개 (3,436개 고유… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-rag-ko.keural-QA-en
keural-synthetic-SFT-en-general
English general-purpose SFT dataset synthetically generated from English Wikipedia using a two-stage pipeline. Unlike RAG-specific datasets, questions and answers contain no references to source documents — the model is expected to answer from its own knowledge, with Wikipedia serving only as a grounding anchor during generation.
Dataset Summary
Item
Value
Total accepted
~1.6M
Total rejected
~8,700
Accept rate
~99.5%
Seed… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/keural-QA-en.
