huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v12-top-only-paper100k-step1500-seed10
0114
liftv12top_only
This repository contains the lift_v12_top_only checkpoint from the matched V12 Paper100K ablation suite for Qwen2.5 7B-to-1.5B knowledge distillation.
Variant
Gap gate with step-level online top-two-layer influence weights only.
Training protocol
- Teacher:
Qwen/Qwen2.5-7B-Instruct - Student initialization:
Qwen/Qwen2.5-1.5B-Instruct - Data:
lift_paper_en_natural_v1/100k(96,000 training examples and a seeded 2,000-example controller split) - Objective: fully on-policy GKD with the variant-specific components above
- Optimizer steps: 1,500
- Global batch size: 64
- Optimizer: AdamW, cosine learning rate from
1e-5to1e-7, weight decay1e-2 - Sampling: temperature
0.9, at most 128 generated tokens - Seed: 10
- Weights: Hugging Face SafeTensors
The full method checkpoint is available at `huggingFacing/qwen25_7B_to_1.5B_v12_onlineif_paper100k_1500`.
Repository ID: huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v12-top-only-paper100k-step1500-seed10.
