CoolFace
Modelpublic

reasoning-degeneration-dev/algo-sft-cellular-automata-distill-qwq

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes7downloads
Model Card

Cellular Automata — QwQ Distillation

LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct fine-tuned on cellular automata via QwQ-32B Distillation.

Part of the Algorithmic SFT vs Distillation experiment studying whether deterministic algorithmic templates teach procedural reasoning more effectively than distillation from large reasoning models.

Training

ParameterValue
Base modelQwen/Qwen2.5-1.5B-Instruct
MethodQwQ-32B Distillation
FrameworkLLaMA-Factory (SFT stage)
LoRA rank64
LoRA targetall linear layers
Learning rate1e-4
Epochs3
Batch size1 (grad accum 16)
Cutoff length32,768 tokens
Training data5,000 QwQ-32B reasoning traces (d5, filtered for correctness). Teacher solve rate: 28.0%

Evaluation (v3, MAX_TOKENS=32768)

SplitAccuracy
Test (in-distribution)40.4%
Harder variant4.8%
Structural OOD22.4%

Notes

Distillation substantially weaker than algorithmic SFT on this domain (40.4% vs 94.6% test).

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "reasoning-degeneration-dev/algo-sft-cellular-automata-distill-qwq")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")

Related Datasets