reasoning-degeneration-dev/algo-sft-long-arithmetic-distill-qwq
06
Long Arithmetic — QwQ Distillation
LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct fine-tuned on long arithmetic via QwQ-32B Distillation.
Part of the Algorithmic SFT vs Distillation experiment studying whether deterministic algorithmic templates teach procedural reasoning more effectively than distillation from large reasoning models.
Training
Evaluation (v3, MAX_TOKENS=32768)
Notes
Nearly tied with algo SFT in-distribution (90.6% vs 92.6%). Slight OOD edge (6.8% vs 0%) but both effectively fail.
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "reasoning-degeneration-dev/algo-sft-long-arithmetic-distill-qwq")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")Related Datasets
- Training data (63K algo traces)
- Distillation data (24K QwQ traces)
- Eval results (aggregate scores)
- Eval questions (11K test/val/harder/OOD)
