CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HAERAE-HUB /KMMLU KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.tabularmultiple-choice100K<n<1M101 likes8.9k downloads3y agoHugging Face02HAERAE-HUB /HAE_RAE_BENCH_1.1The HAE_RAE_BENCH 1.1 is an ongoing project to develop a suite of evaluation tasks designed to test the understanding of models regarding Korean cultural and contextual nuances. Currently, it comprises 13 distinct tasks, with a total of 4900 instances. Please note that although this repository contains datasets from the original HAE-RAE BENCH paper, the contents are not completely identical. Specifically, the reading comprehension subset from the original version has been removed due to… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.1.textmultiple-choice1K<n<10K20 likes3.7k downloads2y agoHugging Face03HAERAE-HUB /KMMLU-HARD KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.textquestion-answering1K<n<10K13 likes3k downloads3y agoHugging Face04HAERAE-HUB /HRM8K | 📖 Paper | 📝 Blog | 🖥️ Code(Coming soon!) | HRM8K We introduce HAE-RAE Math 8K (HRM8K), a bilingual math reasoning benchmark for Korean and English. HRM8K comprises 8,011 instances for evaluation, sourced through a combination of translations from established English benchmarks (e.g., GSM8K, MATH, OmniMath, MMMLU) and original problems curated from existing Korean math exams. Benchmark Overview The HRM8K benchmark consists of two subsets: Korean School Math (KSM):… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRM8K.tabular1K<n<10K24 likes1.7k downloads2y agoHugging Face05HAERAE-HUB /KoSimpleEvaltext100K<n<1M0 likes731 downloads1y agoHugging Face06HAERAE-HUB /KOREAN-WEBTEXT KOREAN-WEBTEXT KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources: cc100 oscar-corpus/OSCAR-2201 oscar-corpus/OSCAR-2109 oscar-corpus/OSCAR-2301 ontocord/CulturaY Additional credible internet sources collected by out team (We are working to add more sources) The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.tabular1M<n<10M49 likes711 downloads2y agoHugging Face07HAERAE-HUB /HAE_RAE_BENCH_1.0The HAE_RAE_BENCH 1.0 is the original implementation of the dataset froom the paper: HAE-RAE BENCH paper. The benchmark is a collection of 1,538 instances across 6 tasks: standard_nomenclature, loan_word, rare_word, general_knowledge, history and reading comprehension. To replicate the studies from the paper, see below. Dataset Overview Task Instances Version Explanation standard_nomenclature 153 v1.0 Multiple-choice questions about Korean standard nomenclatures from… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.0.text1K<n<10K1 likes291 downloads2y agoHugging Face08HAERAE-HUB /NOLLI NOLLI NOLLI (논리) is a benchmark of rule-based logical puzzles in Korean and English, designed with difficulty calibration in mind. It covers 15 task types x 3 difficulty tiers (easy / medium / hard) x 100 examples each, for a total of 7,500 examples (EN 3,000 + KO 4,500). All examples are programmatically generated with solver-verified answers. If you need more data, the generators in the GitHub repository can produce additional examples for any task, language, and difficulty… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/NOLLI.textquestion-answering1K<n<10K9 likes285 downloads2mo agoHugging Face09HAERAE-HUB /csatqa CSAT-QAtabularmultiple-choice1K<n<10K18 likes274 downloads3y agoHugging Face10HAERAE-HUB /KOREAN-SyntheticText-1.5B KOREAN-SyntheticText KOREAN-SyntheticText is a successor of the KOREAN-WEBTEXT project in our mission to create high-quality Korean corpora. The dataset consists of 1.4B tokens generated over 600 H100 hours following the Cosmopedia project. The dataset has been generated using a 100B + open-source LLM fine-tuned on text generation. No filtering has been done yet. text1M<n<10M15 likes244 downloads2y agoHugging Face11HAERAE-HUB /K2-EvalResearch Paper coming soon! K2EvalK^{2} EvalK2Eval K2EvalK^{2} EvalK2Eval is a novel benchmark featuring 90 handwritten instructions that require in-depth knowledge of Korean language and culture for accurate completion. Benchmark Overview The design principle behind K2EvalK^{2} EvalK2Eval centers on collecting instructions that necessitate knowledge specific to Korean culture and context in order to solve. This approach distinguishes our work from simply translating… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/K2-Eval.textn<1K8 likes213 downloads2y agoHugging Face12HAERAE-HUB /KUDGEOfficial data repository for LLM-as-a-Judge & Reward Model: What They Can and Cannot DoTLDR; Automated Evaluators (LLM-as-a-Judge, Reward Models) can be transferred to non-English settings without additional training. (most of the times) Dataset Description At the best of our knowledge, KUDGE is the only, non-English, human-annotated meta-evaluation dataset at this point. Consisted of 5,012 human annotation from native Korean speakers, we expect KUDGE to be widely used as a tool… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KUDGE.tabular1K<n<10K7 likes194 downloads2y agoHugging Face13HAERAE-HUB /K2-FeedbackResearch Paper coming soon! K^2-Feedback K^2-Feedback is a dataset crafted to enhance fine-grained evaluation capabilities in Korean language models. Building upon the Feedback-Collection, K^2-Feedback incorporates instructions specific to Korean culture and linguistics. Dataset Overview K^2-Feedback includes 100,000 samples divided into two distinct subsets: Translated Samples (50,000 entries): This subset consists of samples directly translated from the… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/K2-Feedback.tabular10K<n<100K11 likes92 downloads2y agoHugging Face14HAERAE-HUB /Ko-PIQA Ko-PIQA: Korean Physical Commonsense Reasoning Dataset 📖 Dataset Overview Ko-PIQA is a Korean Physical Commonsense Reasoning dataset designed to complement English-centric benchmarks like PIQA and to include culturally-grounded physical reasoning questions. Total items: 441 Culturally-grounded items: 87 (19.7%)(e.g., kimchi storage, hanbok care, ondol heating) Format: PIQA-style binary choice (solution0 / solution1) Goal: Evaluate Korean LLM physical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/Ko-PIQA.tabularn<1K3 likes90 downloads9mo agoHugging Face15HAERAE-HUB /hret_agent_idavidrein_gpqa_diamond_translatedtabularn<1K0 likes80 downloads2y agoHugging Face16HAERAE-HUB /HAE-RAE-COT-1.5M Dataset Card for "HAE-RAE-COT-1.5M" HAE-RAE-COT-1.5M is a dataset encompassing 1,586,688 samples of questions paired with CoT (Chain of Thought) rationales. The majority of this dataset is a translation of samples from the CoT-Collection, with a portion of samples derived from Korean datasets through the utilization of the gpt-3.5-turbo API. The translation of the CoT-Collection was carried out using the NLLB 600M model. To the best of our knowledge, HAE-RAE-COT-1.5M represents the… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE-RAE-COT-1.5M.text1M<n<10M6 likes72 downloads3y agoHugging Face17HAERAE-HUB /Korean-Human-Judgements※ DISCLAIMER ※ DO NOT USE FOR TRAINING PURPOSES. THE DATA IS FOR EVALUATION ONLY. Korean-Human-Judgements (KHJ) The Korean-Human-Judgements dataset consists of 694 triplets, each containing a question, answer A, and answer B, annotated with human preferences.The original dataset is sourced from three main sources and has undergone a rigorous filtering process to ensure quality and appropriateness. Entries with issues such as sexual content, typographical errors, or unclear… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/Korean-Human-Judgements.textn<1K39 likes71 downloads2y agoHugging Face18HAERAE-HUB /HAERAE-VISION HAERAE-VISION A Korean visual QA benchmark featuring real-world, under-specified questions. Dataset Description This dataset includes two question types: original: Under-specified, authentic user queries explicit: Clarified queries with full context Both share the same images and reference answers, allowing controlled evaluation of query under-specification. Evaluation Code See our GitHub repository for evaluation scripts. Citation… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAERAE-VISION.imagevisual-question-answeringn<1K17 likes62 downloads2mo agoHugging Face19HAERAE-HUB /HR-Instruct-Math-v0.1 Dataset Summary HAERAE-HUB/HR-Instruct-Math-v0.1 is a Math instruction dataset written in the Korean language. This dataset contains evolved instructions aimed at enhancing the learning experience in mathematical concepts. The responses in this dataset are generated from open-source Language Models (LLMs). This is a Proof of Concept (PoC) version, meaning there may be errors or unexpected problems in the dataset. Future iterations will be made to improve the dataset quality.… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HR-Instruct-Math-v0.1.text10K<n<100K10 likes54 downloads2y agoHugging Face20HAERAE-HUB /HRMCR HRMCR HAE-RAE Multi-Step Commonsense Reasoning (HRMCR) is a collection of multi-step reasoning questions automatically generated using templates and algorithms. The questions in HRMCR require LLMs to recall diverse aspects of Korean culture and perform multiple reasoning steps to solve them. 📖 Paper 🖥️ Code (Coming soon!) Example of generated questions in the HRMCR benchmark. The figure showcases generated questions (left) alongside their automatically generated solutions… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRMCR.textn<1K3 likes42 downloads2y agoHugging Face21HAERAE-HUB /QARV-binary-setThe QARV (Question and Answers with Regional Variance) project aims to curate a collection of questions with answers that exhibit regional variations across different nations. text1K<n<10K0 likes37 downloads2y agoHugging Face22bzantium /HAERAE-en HAERAE-en (English-Translated HAERAE-BENCH) This dataset is the English-translated version of the original HAERAE-BENCH, a benchmark designed to evaluate the linguistic and knowledge-based capabilities of Korean language models. For a detailed understanding of the original dataset's construction and motivation, please refer to the paper: HAE-RAE: A New Public Korean-Specific Benchmark Dataset. HAERAE-en was created to enable the evaluation of non-Korean models on the knowledge and… See the full description on the dataset page: https://huggingface.co/datasets/bzantium/HAERAE-en.textquestion-answering1K<n<10K0 likes36 downloads1y agoHugging Face23HAERAE-HUB /kin_20250421text1M<n<10M0 likes35 downloads1y agoHugging Face24HAERAE-HUB /HAE_RAE_BENCH_2.0HAE_RAE_BENCH 2.0 is a miny implementation of Big-Bench consisted of 5 tasks: date_understanding, context_definition_alignment, proverb_unscrambling, 2_digit_multiply, and 3_digit_subtract. Paper Coming Soon (probably). text1K<n<10K4 likes33 downloads2y agoHugging Face25HAERAE-HUB /butterflies_and_moths_vqa Butterflies and Moths VQA Dataset Summary butterflies_and_moths_vqa is a visual question answering (VQA) dataset focused on butterflies and moths. It features tasks such as fine-grained species classification and ecological reasoning. The dataset is designed to benchmark Vision-Language Models (VLMs) for both image-based and text-only training approaches. Key Features Fine-Grained Classification (Type1): Questions requiring detailed species identification.… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/butterflies_and_moths_vqa.imagevisual-question-answeringn<1K0 likes24 downloads2y agoHugging Face26HAERAE-HUB /QARV-preview QARV (Question and Answers with Regional Variance) The QARV (Question and Answers with Regional Variance) project aims to curate a collection of questions with answers that exhibit regional variations across different nations. Version This version contains 1k questions. We are working to add answers for US & Korea. If you are interested in collaborating let us know. text1K<n<10K1 likes20 downloads2y agoHugging Face27HAERAE-HUB /KHJ-RB-Formattextn<1K5 likes9 downloads2y agoHugging Face28dilab-cau /haerae-query-context-stress-v2-extreme HAE-RAE Query/Context Label-Preserving Stress v2 Extreme This repository packages an extreme paired Korean boundary-stress dataset built from HAERAE-HUB/HAE_RAE_BENCH_1.1. What it contains Each row preserves: the original answer options the original gold answer and modifies only the query/context side to make the surface form more tokenization-fragile while keeping: identical non-space character sequence identical Kiwi token signature (form, tag) increased… See the full description on the dataset page: https://huggingface.co/datasets/dilab-cau/haerae-query-context-stress-v2-extreme.tabularmultiple-choicen<1K0 likes9 downloads4mo agoHugging Face29OccasionallyNLP /haerae_testtext1K<n<10K0 likes8 downloads1y agoHugging Face30dilab-cau /haerae-query-context-stress-v3 HAE-RAE Query/Context Label-Preserving Stress v3 This repository packages a v3 paired Korean boundary-stress dataset built from HAERAE-HUB/HAE_RAE_BENCH_1.1. What it contains Each row preserves: the original answer options the original gold answer and modifies only the query/context side to make the surface form more tokenization-fragile while keeping: identical non-space character sequence identical Kiwi token signature (form, tag) increased decoder-tokenizer boundary… See the full description on the dataset page: https://huggingface.co/datasets/dilab-cau/haerae-query-context-stress-v3.tabularmultiple-choice1K<n<10K0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.