tomyimkc/repro-provable-benefits-of-rlvr-over-sft-for-reasoning-models-learning-to-backtrack-ef
0
Reproduction: Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
An open experiment logbook, published with Trackio.
