k-reap
Datasets
All datasets matching “k-reap”K-EXAONE-236B-REAP-calibration-mix
K-EXAONE-236B REAP/NVFP4 Calibration Mix
LGAI-EXAONE/K-EXAONE-236B-A23B의 expert pruning(REAP)과 NVFP4 양자화 calibration을 위해 제작한 믹스.
총 16,780 샘플 / 101,157,434 토큰 (K-EXAONE 토크나이저 기준).
제작 목적
MoE 모델을 one-shot pruning/양자화하면 reasoning 무한 반복(한국어/영어 공통)이 발생하는 문제가 있어,
이를 방지하기 위해 아래 원칙으로 설계:
Context length 다각화: 16 토큰 ~ 245K 토큰 (짧은 지시 → 32K agentic 궤적 → 128K 장문 → 245K needle 스트레스)
한국어 대량 포함 (instruction/reasoning/tool-calling) — K-EXAONE 특화 expert 보호
reasoning trace 원형 보존 —… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/K-EXAONE-236B-REAP-calibration-mix.pose-data-1pose-data-2
