CoolFace
Apppublic

Deepakkambala/repro-provable-benefits-of-rlvr-over-sft-for-reasoning-models-learning-to-backtrack-efficiently

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
3 commits on main
07c3f6a2mo ago

Update logbook: Reproduction: Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently

Deepakkambala
2f3594e2mo ago

Update logbook: Reproduction: Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently

Deepakkambala
c2c01d52mo ago

initial commit

Deepakkambala