datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KMMLU-Pro
KMMLU-Pro
📄 Paper | 🖥️ Code
We introduce KMMLU-PRO, a challenging new benchmark comprising 2,822 problems from the official exams for Korean National Professional Licensure (KNPL), representing highly specialized professions in Korea. To bridge the gap between LLM performance and real-world applicability, our evaluation closely mirrors the official certification criteria including reporting the number of professional licenses each LLM could pass according to these standards. All… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/KMMLU-Pro.FinTexTSKMMLU-Redux
KMMLU-Redux
📄 Paper
We introduce KMMLU-Redux, a reconstructed version of the existing KMMLU, comprising 2,587 problems from Korean National Technical Qualification (KNTQ) exams. We identified several critical issues in the KMMLU, including leaked answers, lack of clarity, ill-posed questions, notation errors, and contamination risks. To address these problems, we conducted rigorous manual examination and decontaminated the dataset from the pre-training corpus to prevent potential… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/KMMLU-Redux.MANTA-1M
Abstract
We introduce MANTA, an automated pipeline that generates high-quality large-scale instruction fine-tuning datasets from massive web corpora while preserving their diversity and scalability. By extracting structured syllabi from web documents and leveraging high-performance LLMs, our approach enables highly effective query-response generation with minimal human intervention. Extensive experiments on 8B-scale LLMs demonstrate that fine-tuning on the MANTA-1M dataset… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/MANTA-1M.KoMT-Bench
KoMT-Bench
Introduction
We present KoMT-Bench, a benchmark designed to evaluate the capability of language models in following instructions in Korean.
KoMT-Bench is an in-house dataset created by translating MT-Bench [1] dataset into Korean and modifying some questions to reflect the characteristics and cultural nuances of the Korean language.
After the initial translation and modification, we requested expert linguists to conduct a thorough review of our benchmark… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/KoMT-Bench.Ko-LongRAG
Abstract
The rapid advancement of large language models (LLMs) significantly enhances long-context Retrieval-Augmented Generation (RAG), yet existing benchmarks focus primarily on English. This leaves low-resource languages without comprehensive evaluation frameworks, limiting their progress in retrieval-based tasks. To bridge this gap, we introduce Ko-LongRAG, the first Korean long-context RAG benchmark. Unlike conventional benchmarks that depend on external retrievers… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/Ko-LongRAG.K-EXAONE-236B-REAP-calibration-mix
K-EXAONE-236B REAP/NVFP4 Calibration Mix
LGAI-EXAONE/K-EXAONE-236B-A23B의 expert pruning(REAP)과 NVFP4 양자화 calibration을 위해 제작한 믹스.
총 16,780 샘플 / 101,157,434 토큰 (K-EXAONE 토크나이저 기준).
제작 목적
MoE 모델을 one-shot pruning/양자화하면 reasoning 무한 반복(한국어/영어 공통)이 발생하는 문제가 있어,
이를 방지하기 위해 아래 원칙으로 설계:
Context length 다각화: 16 토큰 ~ 245K 토큰 (짧은 지시 → 32K agentic 궤적 → 128K 장문 → 245K needle 스트레스)
한국어 대량 포함 (instruction/reasoning/tool-calling) — K-EXAONE 특화 expert 보호
reasoning trace 원형 보존 —… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/K-EXAONE-236B-REAP-calibration-mix.EXAONE-4.0-1.2B-Quantization-MMLUbigbench_mistake_eval_z_EXAONE-Deep-32BLGAI-EXAONE__EXAONE-3.5-7.8B-Instruct-details
Dataset Card for Evaluation run of LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct
Dataset automatically created during the evaluation run of model LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LGAI-EXAONE__EXAONE-3.5-7.8B-Instruct-details._judged_science_traces_original_EXAONE-Deep-32Bko-ko-math-500-test-EXAONE-4.0-1.2B-bonLGAI-EXAONE__EXAONE-3.5-32B-Instruct-details
Dataset Card for Evaluation run of LGAI-EXAONE/EXAONE-3.5-32B-Instruct
Dataset automatically created during the evaluation run of model LGAI-EXAONE/EXAONE-3.5-32B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LGAI-EXAONE__EXAONE-3.5-32B-Instruct-details.math_traces_original_EXAONE-Deep-32Bexp_rob_dfiltered_EXAONE-Deep-32B_2_mbenign_complete_step_t30exp_rob_dfiltered_EXAONE-Deep-32B_2_madversarial_insert_wrong_fact_t10LGAI-EXAONE__EXAONE-3.0-7.8B-Instruct-details
Dataset Card for Evaluation run of LGAI-EXAONE/EXAONE-3.0-7.8B-Instruct
Dataset automatically created during the evaluation run of model LGAI-EXAONE/EXAONE-3.0-7.8B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LGAI-EXAONE__EXAONE-3.0-7.8B-Instruct-details.c_dfiltered_EXAONE-Deep-32B_2_madversarial_insert_wrong_fact_t30j_ablation_force_doubt_logic_EXAONE_Deep_32BEXAONE-32b-HRM8K-AIME-resultscience_traces_original_EXAONE-Deep-32Bc_dfiltered_EXAONE-Deep-32B_2_madversarial_cont_wrong_reasoning_t90c_dfiltered_EXAONE-Deep-32B_2_mbenign_rewrite_trace_t50EXAONE-deep-32B-resultc_dfiltered_science_EXAONE-Deep-32B_mbenign_rewrite_trace_t10exp_rob_dfiltered_EXAONE-Deep-32B_2_mneutral_insert_random_characters_t90c_dfiltered_logic_EXAONE-Deep-32B_mneutral_insert_random_characters_t50preprocessed-ko-ko-math-500-test-EXAONE-4.0-1.2B-bonexp_rob_dfiltered_logic_EXAONE-Deep-32B_madversarial_continue_unrelated_t90exp_rob_dfiltered_logic_EXAONE-Deep-32B_mneutral_add_random_text_t30
