CoolFace
Modelpublic

reasoning-degeneration-dev/algo-sft-long-arithmetic-standard

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes8downloads
Model Card

Long Arithmetic — Standard

LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct fine-tuned on long arithmetic via Algorithmic Template SFT.

Part of the Algorithmic SFT vs Distillation experiment studying whether deterministic algorithmic templates teach procedural reasoning more effectively than distillation from large reasoning models.

Training

ParameterValue
Base modelQwen/Qwen2.5-1.5B-Instruct
MethodAlgorithmic Template SFT
FrameworkLLaMA-Factory (SFT stage)
LoRA rank64
LoRA targetall linear layers
Learning rate1e-4
Epochs3
Batch size4 (grad accum 4)
Cutoff length32,768 tokens
Training data5,000 deterministic carry-propagation traces (d4: 3-digit x 2-3 digit multiply)

Evaluation (v3, MAX_TOKENS=32768)

SplitAccuracy
Test (in-distribution)92.6%
Harder variant21.2%
Structural OOD0.0% (chain operations)

Notes

Strong in-distribution but collapses on structural OOD (chain operations). Neither algo nor distill generalizes here.

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "reasoning-degeneration-dev/algo-sft-long-arithmetic-standard")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")

Related Datasets