CoolFace
Modelpublic

reasoning-degeneration-dev/algo-sft-conlang-morphology-distill-qwq

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes10downloads
Model Card

Conlang Morphology — QwQ Distillation

LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct fine-tuned on conlang morphology via QwQ-32B Distillation.

Part of the Algorithmic SFT vs Distillation experiment studying whether deterministic algorithmic templates teach procedural reasoning more effectively than distillation from large reasoning models.

Training

ParameterValue
Base modelQwen/Qwen2.5-1.5B-Instruct
MethodQwQ-32B Distillation
FrameworkLLaMA-Factory (SFT stage)
LoRA rank64
LoRA targetall linear layers
Learning rate1e-4
Epochs3
Batch size1 (grad accum 16)
Cutoff length32,768 tokens
Training data5,000 QwQ-32B reasoning traces (d7+d5, 3 samples/question, filtered). Teacher solve rate: 15.8-44.3%

Evaluation (v3, MAX_TOKENS=32768)

SplitAccuracy
Test (in-distribution)40.4%
Harder variant11.0%
Structural OOD38.4%

Notes

Large gap vs algo SFT (40.4% vs 98.6%). Distillation traces don't transfer the rule composition procedure.

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "reasoning-degeneration-dev/algo-sft-conlang-morphology-distill-qwq")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")

Related Datasets