CoolFace
Modelpublic

huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step1500

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes18downloads
Model Card

Qwen2.5 7B to 1.5B LiftKD V8 Bilingual 100K - Epoch 2

This is the cumulative epoch-2 checkpoint of a Qwen2.5-1.5B-Instruct student distilled from Qwen2.5-7B-Instruct.

  • —Method: LiftKD V8, fully on-policy GKD JSD with normalized gap gate
  • —Data: 100K bilingual English/Chinese instruction mixture, including 18.75% mathematics
  • —Sequence limits: 384 prompt tokens, 512 total tokens, 128 generated tokens
  • —Precision: BF16 full-parameter training with DeepSpeed ZeRO-2
  • —Global batch size: 64
  • —Seed: 10

The training mixture was internally deduplicated. Evaluation-set decontamination was not performed.