hyper-accel/ci-2layer-llama2-7b
131k
V2: continued KD fine-tune at seq_len 1024, 500 steps, lr 1e-4 from V1 (alpaca-cleaned)
KD-distilled 2-layer student against Llama-2-7B teacher (alpaca-cleaned, 1500 steps, T=2.0, KL loss)
Keep original layers [0, 31] of NousResearch/Llama-2-7b-hf
Add first 2-layer slice of NousResearch/Llama-2-7b-hf
initial commit
