CoolFace
Modelpublic

n-deshpande/yoda-run3-sft

sourceHugging Faceupdated 7d agoView on Hugging Face
0likes22downloads
Model Card

yoda-run3-sft

LoRA adapter (r=32, alpha=64, all linear layers) on Qwen2.5-3B-Instruct that speaks in Yoda's voice while keeping GSM8K accuracy (0.81 on 200 test problems, HF greedy; base 0.855). Trained on the base model's own verifier-correct GSM8K solutions rewritten sentence by sentence into Yoda's voice (DeepSeek V4 Flash) plus Yoda-voiced norobots chat; lr 5e-5, 2 epochs. Top level: the adapter. `merged/`: the adapter merged into the base weights (bf16), the starting point and KL reference of the RLAIF runs. Code, data pipeline and evaluation: https://github.com/harvard-cs2881f26/hw1-wu-ge-deshpande (checkpoint1). Data: n-deshpande/yoda-sft-data.