datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KoHRM-Text-1.4B-prepared-data
KoHRM-Text-1.4B Prepared Data
This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B.
The data is intended for continued pretraining and staged training with the project code at:
https://github.com/LLM-OS-Models/KoHRM-text
https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B
https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K
The upstream architecture and training method are based on:
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.korean-embedding-performance-v1-performance-1m
Korean Embedding Performance v1 — 1M
Qwen3-Embedding-8B의 한국어 retrieval data-scale 실험을 위한 정확히
1,000,000-row 연구·비상업 contrastive dataset이다. release_eligible: false, 통합
라이선스 other이며 upstream source 조건을 재허가하지 않는다.
구성
계열
Rows
비율
역할
nlpai-lab/ko-triplet-v1.0
600,254
60.03%
넓은 한국어 QA/retrieval core
F2 Korean QA/instruction
287,000
28.70%
webfaq, mqa, koalpaca, realQA, komagpie
F2 retrieval task train-family
4,146
0.41%
MIRACL, MrTidy, MLDR
F2 PAWS-X… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-performance-1m.korean-embedding-performance-v1-ablation-200k
Korean Embedding Performance v1 — Ablation 200K
Qwen3-Embedding-8B의 한국어 retrieval continued fine-tuning에서 LoRA/DoRA/부분 및
full fine-tuning, loss, hard-negative 전략을 비교하기 위한 200,000-row 연구·비상업
성능 데이터다. release_eligible: false이며 통합 라이선스는 other다. upstream
source별 조건을 재허가하지 않는다.
구성
계열
Rows
역할
nlpai-lab/ko-triplet-v1.0@1f5d72d
100,254
넓은 한국어 QA/retrieval core
F2 Korean QA/instruction
68,000
webfaq, mqa, koalpaca, realQA, komagpie
F2 retrieval task… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-ablation-200k.korean-embedding-performance-v1-pilot-50k
Korean Embedding Performance v1 — Pilot 50K
주의: 이 revision은 공개 benchmark 성능 후보 학습에 사용하면 안 된다.
사후 15-task exact text-hash 감사에서 평가 query 고유 hash 4개가 확인됐다.
파이프라인·최적화 진단과 contamination ablation에만 남기며, 교체본은
ablation-200k이다.
Qwen3-Embedding 계열의 한국어 retrieval 성능 실험을 위한 50,000-row 연구용
contrastive dataset이다. 각 row는 instruction-aware query, positive passage 1개,
hard/easy negative passage 1–7개를 ms-swift embedding message schema로 저장한다.
사용 조건과 공개 범위
이 저장소의 통합 라이선스는 other다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-pilot-50k.korean-embedding-performance-v1-sionic-retrieval-train-family-4146
Korean Sionic Retrieval Train-Family 4,146
F2LLM-v2가 공개한 Korean MIRACL, MrTidy, MLDR train-family row만 1M
decontaminated curriculum에서 lossless 추출한 target-adaptation dataset이다. 공개
evaluation query는 포함하지 않으며 current-student HN7 mining 전의 source artifact다.
구성과 목적
source
rows
역할
f2_miracl_ko_train
700
MIRACL Korean retrieval train-family
f2_mrtidy_korean_train
1,200
MrTidy Korean train
f2_mldr_ko_train
2,246
MLDR Korean long-document train-family
합계
4… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-retrieval-train-family-4146.korean-legal-retrieval-source-native-250k
Korean Legal Retrieval Source-Native 250K
Legalize-KR의 법령·행정규칙·판례·자치법규 구조에서 query/positive 관계를 추출한
250,000-row 한국어 retrieval dataset이다. release_eligible: false인
target-adapted 연구·비상업 성능 shard이며 통합 라이선스는 other다.
구성과 고정 revision
Source
Revision
Rows
구조 관계
legalize-kr/legalize-kr
db3cd760c14042ee04fd9166e1bdbb662fc999bc
50,000
법령명+조문 → 조문 본문
legalize-kr/admrule-kr
64a5a272909ab5bc077b0ad9519ef31de8febb46
50,000
규칙명+조문 → 조문 본문
legalize-kr/precedent-kr… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-legal-retrieval-source-native-250k.korean-embedding-performance-v1-sionic-squad-train-60k
Korean Embedding — Sionic SQuAD train-family 60K
KorQuAD v1.0의 원본 train split만 질문→정답 문맥 retrieval 형식으로 변환한
60,000-row target-adaptation 데이터다. Sionic retrieval 9종 중
SQuADKorV1의 train-family 신호를 명시적으로 보강한다.
사용 조건과 점수 공개 방식
release_eligible: false인 performance/non-commercial 실험용 composite다. 이
저장소의 통합 라이선스는 other이며 upstream 권리를 재허가하지 않는다. Hub metadata는
KorQuAD source를 CC-BY-ND-4.0으로 표시하고, upstream dataset card 본문은
CC BY-ND 2.0 KR도 명시한다. 사용자는 원 source 조건을 직접 확인해야 한다.
이… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-squad-train-60k.korean-embedding-performance-v1-sionic-autorag-100k
Korean Embedding — Sionic AutoRAG domain 100K
AutoRAG의 금융·상거래·법률 domain retrieval을 보강하기 위한 100,000-row
performance dataset이다. F2LLM-v2 collection의 영어 FIQA/Amazon/Banking77과 중국어
e-commerce/legal QA를 query/positive/negative contrastive schema로 묶었다.
사용 조건과 평가 노출
release_eligible: false인 performance/non-commercial 연구용 composite다. 통합
라이선스는 other이며 F2 collection의 Apache-2.0 표기가 개별 upstream 권리를
재허가하지 않는다.
AutoRAG evaluation repository, query, qrel, corpus는 loader 입력으로… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-autorag-100k.korean-embedding-performance-v1-sionic-health-100k
Korean Embedding — Sionic health multilingual 100K
Qwen3-Embedding 계열의 한국어 PublicHealthQA와 multilingual medical retrieval을
보강하기 위한 100,000-row performance dataset이다. F2LLM-v2 collection의 영어 중심
medical QA/instruction/flashcard와 소량 중국어 WebMedQA를 query/positive/negative
contrastive schema로 묶었다.
사용 조건
release_eligible: false인 performance/non-commercial 연구용 composite다. 통합
라이선스 표기는 other이며 collection card의 Apache-2.0 표기가 각 upstream source의
권리·개인정보·의료 데이터 조건을 재허가하지 않는다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-health-100k.korean-embedding-ko-triplet-hn-pilot-10k
Korean Embedding Ko-Triplet Hard-Negative Pilot 10K
nlpai-lab/ko-triplet-v1.0에서 결정론적으로 뽑은 한국어 retrieval train 10,000행과
validation 512행에 Qwen3-Embedding-8B dense hard negative 4개씩을 붙인 연구용
ms-swift embedding dataset이다.
사용 조건
원 source 카드에 명시적 라이선스가 없어 통합 라이선스는 other, manifest의
release_eligible은 false다. 연구·비상업 성능 실험용이며 이 카드가 원 source의
권리를 재허가하지 않는다.
출처와 sampling
source: nlpai-lab/ko-triplet-v1.0
pinned revision: 1f5d72d21ae8309b5221a588b13930b423385bff… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-ko-triplet-hn-pilot-10k.LFM2.5-KO-Dataset-Index-and-Sources
LFM2.5-KO-Dataset-Index-and-Sources
Snapshot of gyunggyung/LLM-Ko-Datasets README/LICENSE used as a dataset index while building the SFT mix.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-Dataset-Index-and-Sources.ko-legal-embedding-training-v1
Korean Public Legal Embedding Training v1
실제 embedding 학습 queue가 소비하는 provenance-preserving 파생 JSONL이다.
source-native query/positive 관계와 deterministic bootstrap negative로 컴파일했다. 최종 학습 전 current-student hard-negative mining을 수행해야 한다.
rows: 250,000
release eligible: true
visibility: public
use: public redistribution and model training
exact benchmark query/evaluation-text matches: 0
exact retrieval-corpus matches: 0 unique hashes
Sionic 9 및 MTEB task-family train source가 포함될 수… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models2/ko-legal-embedding-training-v1.LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627
LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627
Full Korean CPT mix converted to LFM-style text JSONL, about 4B-token training source.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627.WithinUs-KO-SFT
WithinUs KO SFT Distillation
Korean-translated WithinUs distillation dataset for training Fabliq-KO (LFM2.5-8B-A1B base).
Source / Citation
Original dataset: withinus_mythos_distilled_25k (mathematical reasoning category, 135 rows selected for LFM-SFT format).
This Korean translation is derived from the preprocessed LFM-SFT version. If you use this dataset, please cite both the original source and this Korean-translated version from LLM-OS-Models.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/WithinUs-KO-SFT.Helio-KO-SFT
Helio KO SFT Distillation
Korean-translated Helio deep-reasoning distillation dataset for training Fabliq-KO (LFM2.5-8B-A1B base).
Source / Citation
Original dataset: helio_fable5_distill_reasoning_462x (146 rows of deep-reasoning traces covering security audits, mathematical proofs, biomedical analyses, philosophical treatises, and more).
This Korean translation is derived from the preprocessed LFM-SFT version. If you use this dataset, please cite both the… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Helio-KO-SFT.
