huggingFacing/qwen2.5-7b-to-1.5b-gkd-paper100k-step1500-seed10
0145
vanilla_gkd
This repository contains the vanilla_gkd checkpoint from the matched V12 Paper100K ablation suite for Qwen2.5 7B-to-1.5B knowledge distillation.
Variant
Fully on-policy vanilla GKD baseline.
Training protocol
- Teacher:
Qwen/Qwen2.5-7B-Instruct - Student initialization:
Qwen/Qwen2.5-1.5B-Instruct - Data:
lift_paper_en_natural_v1/100k(96,000 training examples and a seeded 2,000-example controller split) - Objective: fully on-policy GKD with the variant-specific components above
- Optimizer steps: 1,500
- Global batch size: 64
- Optimizer: AdamW, cosine learning rate from
1e-5to1e-7, weight decay1e-2 - Sampling: temperature
0.9, at most 128 generated tokens - Seed: 10
- Weights: Hugging Face SafeTensors
The full method checkpoint is available at `huggingFacing/qwen25_7B_to_1.5B_v12_onlineif_paper100k_1500`.
Repository ID: huggingFacing/qwen2.5-7b-to-1.5b-gkd-paper100k-step1500-seed10.
