CoolFace
7 results

haerae

HAERAE-HUB /KMMLU KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.tabularmultiple-choice100K<n<1M101 likes8.9k downloads3y agoHugging FaceHAERAE-HUB /HAE_RAE_BENCH_1.1The HAE_RAE_BENCH 1.1 is an ongoing project to develop a suite of evaluation tasks designed to test the understanding of models regarding Korean cultural and contextual nuances. Currently, it comprises 13 distinct tasks, with a total of 4900 instances. Please note that although this repository contains datasets from the original HAE-RAE BENCH paper, the contents are not completely identical. Specifically, the reading comprehension subset from the original version has been removed due to… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.1.textmultiple-choice1K<n<10K20 likes3.7k downloads2y agoHugging FaceHAERAE-HUB /KMMLU-HARD KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.textquestion-answering1K<n<10K13 likes3k downloads3y agoHugging FaceHAERAE-HUB /HRM8K | 📖 Paper | 📝 Blog | 🖥️ Code(Coming soon!) | HRM8K We introduce HAE-RAE Math 8K (HRM8K), a bilingual math reasoning benchmark for Korean and English. HRM8K comprises 8,011 instances for evaluation, sourced through a combination of translations from established English benchmarks (e.g., GSM8K, MATH, OmniMath, MMMLU) and original problems curated from existing Korean math exams. Benchmark Overview The HRM8K benchmark consists of two subsets: Korean School Math (KSM):… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRM8K.tabular1K<n<10K24 likes1.7k downloads2y agoHugging FaceHAERAE-HUB /KMMMU KMMMU (Korean MMMU) technical report https://arxiv.org/abs/2604.13058 link to evaluation tutorial! https://github.com/HAE-RAE/KMMMU KMMMU is a Korean version of MMMU: a multimodal benchmark designed to evaluate college-/exam-level reasoning that requires combining images + Korean text. This dataset contains 3,466 questions collected from Korean exam sources including: Civil service recruitment exams National Technical Qualifications National Competency Standard (NCS) exams… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMMU.image1K<n<10K13 likes792 downloads5mo agoHugging FaceHAERAE-HUB /KoSimpleEvaltext100K<n<1M0 likes731 downloads1y agoHugging Face