CoolFace
14 results

kmmlu

HAERAE-HUB /KMMLU KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.tabularmultiple-choice100K<n<1M101 likes9.6k downloads3y agoHugging FaceSabaPivot /KMMLU-Summarized-Chain_of_Thought Dataset Card for Condensed Chain-of-Thought KMMLU Dataset This dataset card provides detailed information about the condensed KMMLU dataset. The dataset has been summarized using Upstage's LLM: Solar-Pro to condense the original KMMLU training and development data while preserving its quality and usability. Additionally, a new column, 'chain_of_thought', has been introduced to align with the reasoning approach outlined in the paper "Chain-of-Thought Prompting Elicits Reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/SabaPivot/KMMLU-Summarized-Chain_of_Thought.tabular100K<n<1M1 likes3.5k downloads2y agoHugging FaceHAERAE-HUB /KMMLU-HARD KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.textquestion-answering1K<n<10K13 likes3.3k downloads3y agoHugging FaceLGAI-EXAONE /KMMLU-Progated KMMLU-Pro 📄 Paper | 🖥️ Code We introduce KMMLU-PRO, a challenging new benchmark comprising 2,822 problems from the official exams for Korean National Professional Licensure (KNPL), representing highly specialized professions in Korea. To bridge the gap between LLM performance and real-world applicability, our evaluation closely mirrors the official certification criteria including reporting the number of professional licenses each LLM could pass according to these standards. All… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/KMMLU-Pro.tabular1K<n<10K33 likes879 downloads1y agoHugging FaceLGAI-EXAONE /KMMLU-Reduxgated KMMLU-Redux 📄 Paper We introduce KMMLU-Redux, a reconstructed version of the existing KMMLU, comprising 2,587 problems from Korean National Technical Qualification (KNTQ) exams. We identified several critical issues in the KMMLU, including leaked answers, lack of clarity, ill-posed questions, notation errors, and contamination risks. To address these problems, we conducted rigorous manual examination and decontaminated the dataset from the pre-training corpus to prevent potential… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/KMMLU-Redux.text1K<n<10K26 likes261 downloads1y agoHugging Facebzantium /KMMLU-en KMMLU-en (English-Translated KMMLU) This dataset is the English-translated version of the original KMMLU (Korean-MMLU) dataset. Description The original KMMLU is a challenging benchmark designed to measure massive multitask language understanding in Korean. It consists of 35,030 expert-level multiple-choice questions across 45 diverse subjects, collected from original Korean exams to capture unique linguistic and cultural nuances. KMMLU-en was created to enable evaluation… See the full description on the dataset page: https://huggingface.co/datasets/bzantium/KMMLU-en.tabularmultiple-choice100K<n<1M0 likes154 downloads1y agoHugging Face