datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KMMLU
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.KMMLU-HARD
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.HRM8K
| 📖 Paper | 📝 Blog | 🖥️ Code(Coming soon!) |
HRM8K
We introduce HAE-RAE Math 8K (HRM8K), a bilingual math reasoning benchmark for Korean and English.
HRM8K comprises 8,011 instances for evaluation, sourced through a combination of translations from established English benchmarks (e.g., GSM8K, MATH, OmniMath, MMMLU) and original problems curated from existing Korean math exams.
Benchmark Overview
The HRM8K benchmark consists of two subsets:
Korean School Math (KSM):… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRM8K.KoSimpleEvalHAE_RAE_BENCH_1.0The HAE_RAE_BENCH 1.0 is the original implementation of the dataset froom the paper: HAE-RAE BENCH paper.
The benchmark is a collection of 1,538 instances across 6 tasks: standard_nomenclature, loan_word, rare_word, general_knowledge, history and reading comprehension.
To replicate the studies from the paper, see below.
Dataset Overview
Task
Instances
Version
Explanation
standard_nomenclature
153
v1.0
Multiple-choice questions about Korean standard nomenclatures from… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.0.K2-EvalResearch Paper coming soon!
K2EvalK^{2} EvalK2Eval
K2EvalK^{2} EvalK2Eval is a novel benchmark featuring 90 handwritten instructions that require in-depth knowledge of Korean language and culture for accurate completion.
Benchmark Overview
The design principle behind K2EvalK^{2} EvalK2Eval centers on collecting instructions that necessitate knowledge specific to Korean culture and context in order to solve. This approach distinguishes our work from simply translating… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/K2-Eval.KUDGEOfficial data repository for LLM-as-a-Judge & Reward Model: What They Can and Cannot DoTLDR; Automated Evaluators (LLM-as-a-Judge, Reward Models) can be transferred to non-English settings without additional training. (most of the times)
Dataset Description
At the best of our knowledge, KUDGE is the only, non-English, human-annotated meta-evaluation dataset at this point.
Consisted of 5,012 human annotation from native Korean speakers, we expect KUDGE to be widely used as a tool… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KUDGE.Ko-PIQA
Ko-PIQA: Korean Physical Commonsense Reasoning Dataset
📖 Dataset Overview
Ko-PIQA is a Korean Physical Commonsense Reasoning dataset designed to complement English-centric benchmarks like PIQA and to include culturally-grounded physical reasoning questions.
Total items: 441
Culturally-grounded items: 87 (19.7%)(e.g., kimchi storage, hanbok care, ondol heating)
Format: PIQA-style binary choice (solution0 / solution1)
Goal: Evaluate Korean LLM physical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/Ko-PIQA.HRMCR
HRMCR
HAE-RAE Multi-Step Commonsense Reasoning (HRMCR) is a collection of multi-step reasoning questions automatically generated using templates and algorithms.
The questions in HRMCR require LLMs to recall diverse aspects of Korean culture and perform multiple reasoning steps to solve them.
📖 Paper
🖥️ Code (Coming soon!)
Example of generated questions in the HRMCR benchmark. The figure showcases generated questions (left) alongside
their automatically generated solutions… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRMCR.QARV-binary-setThe QARV (Question and Answers with Regional Variance) project aims to curate a collection of questions with answers that exhibit regional variations across different nations.
HAERAE-en
HAERAE-en (English-Translated HAERAE-BENCH)
This dataset is the English-translated version of the original HAERAE-BENCH, a benchmark designed to evaluate the linguistic and knowledge-based capabilities of Korean language models. For a detailed understanding of the original dataset's construction and motivation, please refer to the paper: HAE-RAE: A New Public Korean-Specific Benchmark Dataset.
HAERAE-en was created to enable the evaluation of non-Korean models on the knowledge and… See the full description on the dataset page: https://huggingface.co/datasets/bzantium/HAERAE-en.kin_20250421HAE_RAE_BENCH_2.0HAE_RAE_BENCH 2.0 is a miny implementation of Big-Bench consisted of 5 tasks: date_understanding, context_definition_alignment, proverb_unscrambling, 2_digit_multiply,
and 3_digit_subtract.
Paper Coming Soon (probably).
