tomyimkc/repro-provable-benefits-of-rlvr-over-sft-for-reasoning-models-learning-to-backtrack-ef
0
Update logbook: Reproduction: Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
initial commit
Update logbook: Reproduction: Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
initial commit