juihuichung/awakening-rl-replay100
0104
awakening-rl-replay100
The RL-base "+100 agentic replay" arm of "Minimal Agentic Replay Recovers Tool Use in Formal-Math Fine-Tuned LLMs" (COLM 2026): awakening-goedel-v2-32b-rl fine-tuned on the same 100 behavior-tilted Lean agentic traces as the SFT arm (128 epochs, lr 1e-5 cosine, checkpoint-128). Recovers BFCL Non-Live to 85.8 ("Goedel-RL-RAG-BT100" in the paper's score tables), showing the minimal-replay recovery holds on the RL checkpoint as well as the SFT one.
Related
- Code, data manifests, training recipes, eval scripts: github.com/juihuichung/awakening
- Training data: juihuichung/awakening-data
- Models: awakening-goedel-v2-32b-sft · awakening-goedel-v2-32b-rl · awakening-sft-replay100 · awakening-rl-replay100
