CoolFace
Modelpublic

reasoning-degeneration-dev/algo-sft-long-arithmetic-chunked

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes11downloads
Model Card

Long Arithmetic — Chunked

LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct fine-tuned on long arithmetic via Algorithmic Template SFT.

Part of the Algorithmic SFT vs Distillation experiment studying whether deterministic algorithmic templates teach procedural reasoning more effectively than distillation from large reasoning models.

Training

ParameterValue
Base modelQwen/Qwen2.5-1.5B-Instruct
MethodAlgorithmic Template SFT
FrameworkLLaMA-Factory (SFT stage)
LoRA rank64
LoRA targetall linear layers
Learning rate1e-4
Epochs3
Batch size4 (grad accum 4)
Cutoff length32,768 tokens
Training data5,000 deterministic chunked multiplication traces (d4)

Evaluation (v3, MAX_TOKENS=32768)

SplitAccuracy
Test (in-distribution)86.2%
Harder variant13.2%
Structural OOD0.0%

Notes

Weaker than standard variant. Same OOD failure.

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "reasoning-degeneration-dev/algo-sft-long-arithmetic-chunked")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")

Related Datasets