CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HAERAE-HUB /KMMLU KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.tabularmultiple-choice100K<n<1M101 likes8.9k downloads3y agoHugging Face02HAERAE-HUB /HRM8K | 📖 Paper | 📝 Blog | 🖥️ Code(Coming soon!) | HRM8K We introduce HAE-RAE Math 8K (HRM8K), a bilingual math reasoning benchmark for Korean and English. HRM8K comprises 8,011 instances for evaluation, sourced through a combination of translations from established English benchmarks (e.g., GSM8K, MATH, OmniMath, MMMLU) and original problems curated from existing Korean math exams. Benchmark Overview The HRM8K benchmark consists of two subsets: Korean School Math (KSM):… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRM8K.tabular1K<n<10K24 likes1.7k downloads2y agoHugging Face03HAERAE-HUB /KOREAN-WEBTEXT KOREAN-WEBTEXT KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources: cc100 oscar-corpus/OSCAR-2201 oscar-corpus/OSCAR-2109 oscar-corpus/OSCAR-2301 ontocord/CulturaY Additional credible internet sources collected by out team (We are working to add more sources) The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.tabular1M<n<10M49 likes711 downloads2y agoHugging Face04HAERAE-HUB /csatqa CSAT-QAtabularmultiple-choice1K<n<10K18 likes274 downloads3y agoHugging Face05HAERAE-HUB /KUDGEOfficial data repository for LLM-as-a-Judge & Reward Model: What They Can and Cannot DoTLDR; Automated Evaluators (LLM-as-a-Judge, Reward Models) can be transferred to non-English settings without additional training. (most of the times) Dataset Description At the best of our knowledge, KUDGE is the only, non-English, human-annotated meta-evaluation dataset at this point. Consisted of 5,012 human annotation from native Korean speakers, we expect KUDGE to be widely used as a tool… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KUDGE.tabular1K<n<10K7 likes194 downloads2y agoHugging Face06HAERAE-HUB /K2-FeedbackResearch Paper coming soon! K^2-Feedback K^2-Feedback is a dataset crafted to enhance fine-grained evaluation capabilities in Korean language models. Building upon the Feedback-Collection, K^2-Feedback incorporates instructions specific to Korean culture and linguistics. Dataset Overview K^2-Feedback includes 100,000 samples divided into two distinct subsets: Translated Samples (50,000 entries): This subset consists of samples directly translated from the… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/K2-Feedback.tabular10K<n<100K11 likes92 downloads2y agoHugging Face07HAERAE-HUB /Ko-PIQA Ko-PIQA: Korean Physical Commonsense Reasoning Dataset 📖 Dataset Overview Ko-PIQA is a Korean Physical Commonsense Reasoning dataset designed to complement English-centric benchmarks like PIQA and to include culturally-grounded physical reasoning questions. Total items: 441 Culturally-grounded items: 87 (19.7%)(e.g., kimchi storage, hanbok care, ondol heating) Format: PIQA-style binary choice (solution0 / solution1) Goal: Evaluate Korean LLM physical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/Ko-PIQA.tabularn<1K3 likes90 downloads9mo agoHugging Face08HAERAE-HUB /hret_agent_idavidrein_gpqa_diamond_translatedtabularn<1K0 likes80 downloads2y agoHugging Face09dilab-cau /haerae-query-context-stress-v2-extreme HAE-RAE Query/Context Label-Preserving Stress v2 Extreme This repository packages an extreme paired Korean boundary-stress dataset built from HAERAE-HUB/HAE_RAE_BENCH_1.1. What it contains Each row preserves: the original answer options the original gold answer and modifies only the query/context side to make the surface form more tokenization-fragile while keeping: identical non-space character sequence identical Kiwi token signature (form, tag) increased… See the full description on the dataset page: https://huggingface.co/datasets/dilab-cau/haerae-query-context-stress-v2-extreme.tabularmultiple-choicen<1K0 likes9 downloads4mo agoHugging Face10dilab-cau /haerae-query-context-stress-v3 HAE-RAE Query/Context Label-Preserving Stress v3 This repository packages a v3 paired Korean boundary-stress dataset built from HAERAE-HUB/HAE_RAE_BENCH_1.1. What it contains Each row preserves: the original answer options the original gold answer and modifies only the query/context side to make the surface form more tokenization-fragile while keeping: identical non-space character sequence identical Kiwi token signature (form, tag) increased decoder-tokenizer boundary… See the full description on the dataset page: https://huggingface.co/datasets/dilab-cau/haerae-query-context-stress-v3.tabularmultiple-choice1K<n<10K0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.