CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LLM-OS-Models /KoHRM-Text-1.4B-prepared-data KoHRM-Text-1.4B Prepared Data This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B. The data is intended for continued pretraining and staged training with the project code at: https://github.com/LLM-OS-Models/KoHRM-text https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K The upstream architecture and training method are based on: Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.tabulartext-generationn<1K1 likes1k downloads4mo agoHugging Face02LLM-OS-Models /korean-embedding-performance-v1-performance-1m Korean Embedding Performance v1 — 1M Qwen3-Embedding-8B의 한국어 retrieval data-scale 실험을 위한 정확히 1,000,000-row 연구·비상업 contrastive dataset이다. release_eligible: false, 통합 라이선스 other이며 upstream source 조건을 재허가하지 않는다. 구성 계열 Rows 비율 역할 nlpai-lab/ko-triplet-v1.0 600,254 60.03% 넓은 한국어 QA/retrieval core F2 Korean QA/instruction 287,000 28.70% webfaq, mqa, koalpaca, realQA, komagpie F2 retrieval task train-family 4,146 0.41% MIRACL, MrTidy, MLDR F2 PAWS-X… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-performance-1m.textsentence-similarity1M<n<10M0 likes101 downloads3mo agoHugging Face03LLM-OS-Models /korean-embedding-performance-v1-ablation-200k Korean Embedding Performance v1 — Ablation 200K Qwen3-Embedding-8B의 한국어 retrieval continued fine-tuning에서 LoRA/DoRA/부분 및 full fine-tuning, loss, hard-negative 전략을 비교하기 위한 200,000-row 연구·비상업 성능 데이터다. release_eligible: false이며 통합 라이선스는 other다. upstream source별 조건을 재허가하지 않는다. 구성 계열 Rows 역할 nlpai-lab/ko-triplet-v1.0@1f5d72d 100,254 넓은 한국어 QA/retrieval core F2 Korean QA/instruction 68,000 webfaq, mqa, koalpaca, realQA, komagpie F2 retrieval task… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-ablation-200k.textsentence-similarity100K<n<1M0 likes95 downloads3mo agoHugging Face04LLM-OS-Models /korean-embedding-performance-v1-pilot-50k Korean Embedding Performance v1 — Pilot 50K 주의: 이 revision은 공개 benchmark 성능 후보 학습에 사용하면 안 된다. 사후 15-task exact text-hash 감사에서 평가 query 고유 hash 4개가 확인됐다. 파이프라인·최적화 진단과 contamination ablation에만 남기며, 교체본은 ablation-200k이다. Qwen3-Embedding 계열의 한국어 retrieval 성능 실험을 위한 50,000-row 연구용 contrastive dataset이다. 각 row는 instruction-aware query, positive passage 1개, hard/easy negative passage 1–7개를 ms-swift embedding message schema로 저장한다. 사용 조건과 공개 범위 이 저장소의 통합 라이선스는 other다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-pilot-50k.texttext-retrieval10K<n<100K0 likes91 downloads3mo agoHugging Face05LLM-OS-Models /korean-embedding-performance-v1-sionic-retrieval-train-family-4146 Korean Sionic Retrieval Train-Family 4,146 F2LLM-v2가 공개한 Korean MIRACL, MrTidy, MLDR train-family row만 1M decontaminated curriculum에서 lossless 추출한 target-adaptation dataset이다. 공개 evaluation query는 포함하지 않으며 current-student HN7 mining 전의 source artifact다. 구성과 목적 source rows 역할 f2_miracl_ko_train 700 MIRACL Korean retrieval train-family f2_mrtidy_korean_train 1,200 MrTidy Korean train f2_mldr_ko_train 2,246 MLDR Korean long-document train-family 합계 4… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-retrieval-train-family-4146.texttext-retrieval1K<n<10K0 likes61 downloads3mo agoHugging Face06LLM-OS-Models /korean-legal-retrieval-source-native-250k Korean Legal Retrieval Source-Native 250K Legalize-KR의 법령·행정규칙·판례·자치법규 구조에서 query/positive 관계를 추출한 250,000-row 한국어 retrieval dataset이다. release_eligible: false인 target-adapted 연구·비상업 성능 shard이며 통합 라이선스는 other다. 구성과 고정 revision Source Revision Rows 구조 관계 legalize-kr/legalize-kr db3cd760c14042ee04fd9166e1bdbb662fc999bc 50,000 법령명+조문 → 조문 본문 legalize-kr/admrule-kr 64a5a272909ab5bc077b0ad9519ef31de8febb46 50,000 규칙명+조문 → 조문 본문 legalize-kr/precedent-kr… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-legal-retrieval-source-native-250k.textsentence-similarity100K<n<1M0 likes51 downloads3mo agoHugging Face07LLM-OS-Models /korean-embedding-performance-v1-sionic-squad-train-60k Korean Embedding — Sionic SQuAD train-family 60K KorQuAD v1.0의 원본 train split만 질문→정답 문맥 retrieval 형식으로 변환한 60,000-row target-adaptation 데이터다. Sionic retrieval 9종 중 SQuADKorV1의 train-family 신호를 명시적으로 보강한다. 사용 조건과 점수 공개 방식 release_eligible: false인 performance/non-commercial 실험용 composite다. 이 저장소의 통합 라이선스는 other이며 upstream 권리를 재허가하지 않는다. Hub metadata는 KorQuAD source를 CC-BY-ND-4.0으로 표시하고, upstream dataset card 본문은 CC BY-ND 2.0 KR도 명시한다. 사용자는 원 source 조건을 직접 확인해야 한다. 이… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-squad-train-60k.textsentence-similarity10K<n<100K0 likes44 downloads3mo agoHugging Face08LLM-OS-Models /korean-embedding-performance-v1-sionic-autorag-100k Korean Embedding — Sionic AutoRAG domain 100K AutoRAG의 금융·상거래·법률 domain retrieval을 보강하기 위한 100,000-row performance dataset이다. F2LLM-v2 collection의 영어 FIQA/Amazon/Banking77과 중국어 e-commerce/legal QA를 query/positive/negative contrastive schema로 묶었다. 사용 조건과 평가 노출 release_eligible: false인 performance/non-commercial 연구용 composite다. 통합 라이선스는 other이며 F2 collection의 Apache-2.0 표기가 개별 upstream 권리를 재허가하지 않는다. AutoRAG evaluation repository, query, qrel, corpus는 loader 입력으로… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-autorag-100k.textsentence-similarity100K<n<1M0 likes44 downloads3mo agoHugging Face09LLM-OS-Models /korean-embedding-performance-v1-sionic-health-100k Korean Embedding — Sionic health multilingual 100K Qwen3-Embedding 계열의 한국어 PublicHealthQA와 multilingual medical retrieval을 보강하기 위한 100,000-row performance dataset이다. F2LLM-v2 collection의 영어 중심 medical QA/instruction/flashcard와 소량 중국어 WebMedQA를 query/positive/negative contrastive schema로 묶었다. 사용 조건 release_eligible: false인 performance/non-commercial 연구용 composite다. 통합 라이선스 표기는 other이며 collection card의 Apache-2.0 표기가 각 upstream source의 권리·개인정보·의료 데이터 조건을 재허가하지 않는다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-health-100k.textsentence-similarity100K<n<1M0 likes34 downloads3mo agoHugging Face10LLM-OS-Models /korean-embedding-ko-triplet-hn-pilot-10k Korean Embedding Ko-Triplet Hard-Negative Pilot 10K nlpai-lab/ko-triplet-v1.0에서 결정론적으로 뽑은 한국어 retrieval train 10,000행과 validation 512행에 Qwen3-Embedding-8B dense hard negative 4개씩을 붙인 연구용 ms-swift embedding dataset이다. 사용 조건 원 source 카드에 명시적 라이선스가 없어 통합 라이선스는 other, manifest의 release_eligible은 false다. 연구·비상업 성능 실험용이며 이 카드가 원 source의 권리를 재허가하지 않는다. 출처와 sampling source: nlpai-lab/ko-triplet-v1.0 pinned revision: 1f5d72d21ae8309b5221a588b13930b423385bff… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-ko-triplet-hn-pilot-10k.texttext-retrieval10K<n<100K0 likes34 downloads3mo agoHugging Face11LLM-OS-Models /LFM2.5-KO-Dataset-Index-and-Sources LFM2.5-KO-Dataset-Index-and-Sources Snapshot of gyunggyung/LLM-Ko-Datasets README/LICENSE used as a dataset index while building the SFT mix. This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow. Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-Dataset-Index-and-Sources.textn<1K0 likes20 downloads3mo agoHugging Face12LLM-OS-Models2 /ko-legal-embedding-training-v1 Korean Public Legal Embedding Training v1 실제 embedding 학습 queue가 소비하는 provenance-preserving 파생 JSONL이다. source-native query/positive 관계와 deterministic bootstrap negative로 컴파일했다. 최종 학습 전 current-student hard-negative mining을 수행해야 한다. rows: 250,000 release eligible: true visibility: public use: public redistribution and model training exact benchmark query/evaluation-text matches: 0 exact retrieval-corpus matches: 0 unique hashes Sionic 9 및 MTEB task-family train source가 포함될 수… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models2/ko-legal-embedding-training-v1.textsentence-similarity100K<n<1M0 likes20 downloads2mo agoHugging Face13LLM-OS-Models /LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627 LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627 Full Korean CPT mix converted to LFM-style text JSONL, about 4B-token training source. This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow. Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627.text1M<n<10M0 likes18 downloads3mo agoHugging Face14LLM-OS-Models /WithinUs-KO-SFT WithinUs KO SFT Distillation Korean-translated WithinUs distillation dataset for training Fabliq-KO (LFM2.5-8B-A1B base). Source / Citation Original dataset: withinus_mythos_distilled_25k (mathematical reasoning category, 135 rows selected for LFM-SFT format). This Korean translation is derived from the preprocessed LFM-SFT version. If you use this dataset, please cite both the original source and this Korean-translated version from LLM-OS-Models.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/WithinUs-KO-SFT.texttext-generationn<1K0 likes12 downloads3mo agoHugging Face15LLM-OS-Models /Helio-KO-SFT Helio KO SFT Distillation Korean-translated Helio deep-reasoning distillation dataset for training Fabliq-KO (LFM2.5-8B-A1B base). Source / Citation Original dataset: helio_fable5_distill_reasoning_462x (146 rows of deep-reasoning traces covering security audits, mathematical proofs, biomedical analyses, philosophical treatises, and more). This Korean translation is derived from the preprocessed LFM-SFT version. If you use this dataset, please cite both the… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Helio-KO-SFT.texttext-generationn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.