datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KorMedMCQA
KorMedMCQA : Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations
We present KorMedMCQA, the first Korean Medical Multiple-Choice Question
Answering benchmark, derived from professional healthcare licensing
examinations conducted in Korea between 2012 and 2024. The dataset contains
7,469 questions from examinations for doctor, nurse, pharmacist, and dentist,
covering a wide range of medical disciplines. We evaluate the performance… See the full description on the dataset page: https://huggingface.co/datasets/sean0042/KorMedMCQA.korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.KoreanSAT
KoreanSAT Benchmark
Current Topics over years
Math
o1-preview
Year
2022
2023
2024
2025
General
62
50
62
54
Probability & Statistics
22
20
26
22
Calculus
18
18
15
19
Geometry
15
22
18
19
General+Prob.
84
70
88
76
General+Calc.
80
68
77
73
General+Geom.
77
72
80
73
Average
80.3
70
81.7
74
korean-current-law-bar-exam-sft-1000
Korean Current-Law Bar Exam SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 스타일 SFT 데이터 1,000문항입니다.
이 데이터셋은 법무부 기출문제를 복제하지 않습니다. 기존 gyung/korean-bar-exam-moj-multiple-choice의 data/questions.csv는 난도와 과목 분포 참고 및 제15회 중복 방지 기준으로만 사용했습니다.
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 과목 분포, 제15회 유사도 QA 결과입니다.
Columns
question_text: 문제와 5개 선택지
answer: 정답 번호, 1부터 5… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-current-law-bar-exam-sft-1000.korean-traditional-liquor-dataset
Dataset Card for Korean Traditional Liquor Corpus (886–1947)
Dataset Summary
A corpus of Korean traditional liquors and brewing culture curated from 132 historical texts written between 886 and 1947. It includes paragraph-level chunks for IR/RAG and document-/process-level metadata for cultural and historical analyses.
Documents: 132 historical texts (886–1947)
Entries: 3,519 total
🍶 Traditional liquors: 2,876
🍞 Fermentation starters (Nuruk): 193
📦 Others… See the full description on the dataset page: https://huggingface.co/datasets/Jaeuk-Han/korean-traditional-liquor-dataset.KORA-Benchmark
KORA Benchmark
Resources for reproducing KORA: Adaptive Multi-Agent Orchestrated Retrieval over Knowledge Graphs — including the BioCQ benchmark dataset, entity resolution indexes, and the combined biomedical knowledge graph.
Repository Contents
Path
Description
benchmark/
BioCQ question splits (train / val / test / full)
indexes/scispacy*/
Pre-built SciSpaCy entity resolution indexes (~1 GB)
indexes/ark_bm25/
Pre-built ARK BM25 retrieval indexes… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-kgqa/KORA-Benchmark.
