haerae
Datasets
All datasets matching “haerae”KMMLU
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.HAE_RAE_BENCH_1.1The HAE_RAE_BENCH 1.1 is an ongoing project to develop a suite of evaluation tasks designed to test the
understanding of models regarding Korean cultural and contextual nuances.
Currently, it comprises 13 distinct tasks, with a total of 4900 instances.
Please note that although this repository contains datasets from the original HAE-RAE BENCH paper,
the contents are not completely identical. Specifically, the reading comprehension subset from the original version has been removed due to… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.1.KMMLU-HARD
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.HRM8K
| 📖 Paper | 📝 Blog | 🖥️ Code(Coming soon!) |
HRM8K
We introduce HAE-RAE Math 8K (HRM8K), a bilingual math reasoning benchmark for Korean and English.
HRM8K comprises 8,011 instances for evaluation, sourced through a combination of translations from established English benchmarks (e.g., GSM8K, MATH, OmniMath, MMMLU) and original problems curated from existing Korean math exams.
Benchmark Overview
The HRM8K benchmark consists of two subsets:
Korean School Math (KSM):… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRM8K.KMMMU
KMMMU (Korean MMMU)
technical report https://arxiv.org/abs/2604.13058
link to evaluation tutorial! https://github.com/HAE-RAE/KMMMU
KMMMU is a Korean version of MMMU: a multimodal benchmark designed to evaluate college-/exam-level reasoning that requires combining images + Korean text.
This dataset contains 3,466 questions collected from Korean exam sources including:
Civil service recruitment exams
National Technical Qualifications
National Competency Standard (NCS) exams… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMMU.KoSimpleEval
