CoolFace
Modelpublic

huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v12-head-only-paper100k-step1500-seed10

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes125downloads
Model Card

liftv12head_only

This repository contains the lift_v12_head_only checkpoint from the matched V12 Paper100K ablation suite for Qwen2.5 7B-to-1.5B knowledge distillation.

Variant

Gap gate with step-level online LM-head influence weights only.

Training protocol

  • —Teacher: Qwen/Qwen2.5-7B-Instruct
  • —Student initialization: Qwen/Qwen2.5-1.5B-Instruct
  • —Data: lift_paper_en_natural_v1/100k (96,000 training examples and a seeded 2,000-example controller split)
  • —Objective: fully on-policy GKD with the variant-specific components above
  • —Optimizer steps: 1,500
  • —Global batch size: 64
  • —Optimizer: AdamW, cosine learning rate from 1e-5 to 1e-7, weight decay 1e-2
  • —Sampling: temperature 0.9, at most 128 generated tokens
  • —Seed: 10
  • —Weights: Hugging Face SafeTensors

The full method checkpoint is available at `huggingFacing/qwen25_7B_to_1.5B_v12_onlineif_paper100k_1500`.

Repository ID: huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v12-head-only-paper100k-step1500-seed10.