CoolFace
Modelpublic

reasoning-degeneration-dev/algo-sft-long-arithmetic-distill-qwq

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes6downloads
Model Card

Long Arithmetic — QwQ Distillation

LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct fine-tuned on long arithmetic via QwQ-32B Distillation.

Part of the Algorithmic SFT vs Distillation experiment studying whether deterministic algorithmic templates teach procedural reasoning more effectively than distillation from large reasoning models.

Training

ParameterValue
Base modelQwen/Qwen2.5-1.5B-Instruct
MethodQwQ-32B Distillation
FrameworkLLaMA-Factory (SFT stage)
LoRA rank64
LoRA targetall linear layers
Learning rate1e-4
Epochs3
Batch size1 (grad accum 16)
Cutoff length32,768 tokens
Training data5,000 QwQ-32B reasoning traces (d4, filtered). Teacher solve rate: 43.8%

Evaluation (v3, MAX_TOKENS=32768)

SplitAccuracy
Test (in-distribution)90.6%
Harder variant8.4%
Structural OOD6.8%

Notes

Nearly tied with algo SFT in-distribution (90.6% vs 92.6%). Slight OOD edge (6.8% vs 0%) but both effectively fail.

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "reasoning-degeneration-dev/algo-sft-long-arithmetic-distill-qwq")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")

Related Datasets