juihuichung/awakening-sft-replay100
0102
awakening-sft-replay100
The headline "+100 agentic replay" arm of "Minimal Agentic Replay Recovers Tool Use in Formal-Math Fine-Tuned LLMs" (COLM 2026): awakening-goedel-v2-32b-sft (BFCL Non-Live 4.6, tool-calling fully collapsed) fine-tuned on just 100 behavior-tilted Lean agentic traces (OMR-sharegpt-P2.behavior_tilted.sample_100, 128 epochs, lr 1e-5 cosine, cutoff 16384, checkpoint-128). Recovers BFCL Non-Live to 83.8 while retaining Lean-4 proving ability.
Related
- Code, data manifests, training recipes, eval scripts: github.com/juihuichung/awakening
- Training data: juihuichung/awakening-data
- Models: awakening-goedel-v2-32b-sft · awakening-goedel-v2-32b-rl · awakening-sft-replay100 · awakening-rl-replay100
