CoolFace
Modelpublic

Baekpica/K-EXAONE-152B-A23B-REAP-80E-BF16-v2

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes22downloads
Model Card

K-EXAONE-152B-A23B-REAP-80E-BF16-v2

English

This is a conservatively named v2 of Baekpica/K-EXAONE-152B-A23B-REAP-80E-BF16. The parent is a REAP-pruned version of LGAI-EXAONE/K-EXAONE-236B-A23B, with 80 routed experts per MoE layer and 23B active parameters.

v2 applies a small post-pruning recovery run. It is not described as fully healed: measured gains are modest, and reasoning length remains sensitive to the generation cap.

Recovery recipe

  1. 1.Broad continued pretraining
  2. 2.523,776 target tokens, sequence length 1,024
  3. 3.Korean, code, math, science, tool use, agentic SWE, long-context, and stress buckets
  4. 4.Reasoning/Korean SFT
  5. 5.50 optimizer steps, 161,073 assistant-loss tokens
  6. 6.complete chats with final end-of-turn tokens
  7. 7.Trainable modules
  8. 8.all 47 routed-expert routers at full rank: 23,101,440 parameters
  9. 9.rank-2 expert LoRA on routed experts in MoE layers 36–47: 47,185,920 parameters
  10. 10.total trainable parameters: 70,287,360
  11. 11.Frozen modules
  12. 12.attention, embeddings, LM head, shared experts, early routed experts, and MTP head

The run used BF16 on four A100 80GB GPUs. Router load-balancing loss and router z-loss were included. No teacher model or distillation was used.

Evaluation

Small smoke test, not a general benchmark:

  • —vLLM 0.24.0, tensor parallel 4, A100 BF16
  • —12 Korean/English reasoning prompts
  • —seeds 42, 123, and 2026
  • —temperature=1.0, top_p=0.95, max_tokens=3072
  • —MTP disabled for the comparison
ModelClosed `</think>`Emitted EOTHit 3,072-token cap
Parent BF1628/3627/369/36
v227/3627/369/36

The held-out assistant CE improved from 1.18853 to 1.18012 and perplexity from 3.28226 to 3.25477. There were zero dead routed experts before and after.

Two recovery candidates were compared. A longer 110-step SFT candidate closed </think> more often, but produced more tokens and higher repeated n-gram ratios. This 50-step candidate was selected as the more conservative trade-off:

CandidateGenerated tokensRepeated 2-gramRepeated 4-gramRepeated 8-gram
110-step77,4840.51660.24050.0732
v2 / 50-step74,8620.51190.23430.0680

The full integrity and generation records are included as integrity_report.json, staged_healing_report.json, and evaluation_vllm024_3072.json.

Usage

transformers >= 5.1.0 is required for native exaone_moe support.

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Baekpica/K-EXAONE-152B-A23B-REAP-80E-BF16-v2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    dtype="bfloat16",
    device_map="auto",
)
bash
vllm serve Baekpica/K-EXAONE-152B-A23B-REAP-80E-BF16-v2 \
  --tensor-parallel-size 4 \
  --reasoning-parser deepseek_v3 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":2}'

Recommended starting sampling parameters are temperature=1.0 and top_p=0.95. Do not interpret a 2K generation cap as proof of non-termination for this model; the smoke tests showed materially more normal endings at a 3K cap.

Limitations

  • —Improvements are small and are not uniform across prompts or seeds.
  • —This model does not restore the behavior of the unpruned 236B teacher.
  • —Long reasoning, Korean reasoning, tool use, and long-context behavior require task-specific evaluation before production use.
  • —The A100 test used ordinary BF16 execution and does not validate every serving stack.

한국어

이 모델은 Baekpica/K-EXAONE-152B-A23B-REAP-80E-BF16에 소규모 post-pruning recovery를 적용한 v2입니다. 완전한 복구를 의미하는 HEALED 명칭은 사용하지 않았습니다.

적용한 절차는 다음과 같습니다.

  • —523,776 target token의 broad CPT
  • —모든 47개 router full-rank 학습
  • —MoE 36–47 layer의 routed expert에 rank-2 LoRA
  • —50-step Reasoning/한국어 SFT, 161,073 assistant-loss token
  • —Teacher 및 distillation 미사용

vLLM 0.24.0, 4×A100, 3개 seed, 36회 생성의 3,072-token 비교에서 부모와 v2는 EOT 27/36, cap 도달 9/36으로 동률이었습니다. Holdout CE는 1.18853에서 1.18012로 개선됐습니다. 더 긴 110-step 후보보다 반복 n-gram과 전체 생성량이 낮아 50-step 후보를 v2로 선택했습니다.

개선 폭은 제한적이며 모든 prompt에서 일관된 향상을 보장하지 않습니다. 특히 reasoning 길이는 max_tokens에 민감하므로 2K 제한만으로 무한 reasoning 여부를 판단하지 않는 것을 권장합니다. 실제 배포 전 한국어, reasoning, tool use 및 long-context 태스크를 별도로 검증하세요.

License

This model is distributed under the same K-EXAONE AI Model License as the original model. See LICENSE for the complete terms.