BayesRL/Llama3.1-IVON-SFT-8B
Qwen2.5Math-IVON-SFT-7B
📄 Paper: Parameter Exploration for RLVR via Variational Learning · arXiv:2608.09805
📦 Code: insait-institute/c3po
Qwen2.5-Math 7B supervised-fine-tuned with the variational optimizer IVON, from the paper "Parameter Exploration for RLVR via Variational Learning".
This is a warm-start checkpoint: SFT'ing with IVON yields not just point weights but an approximate Gaussian posterior over them (a mean and a diagonal Hessian/precision estimate). That posterior is the learned prior used to seed the 3PO RLVR runs (B3PO / M3PO / C3PO), where weight perturbations sampled from it drive parameter-space exploration.
Training
Usage
Loads as a standard causal LM:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("BayesRL/Qwen2.5Math-IVON-SFT-7B")
tok = AutoTokenizer.from_pretrained("BayesRL/Qwen2.5Math-IVON-SFT-7B")To use it as the warm-start prior for 3PO RLVR, load the IVON optimizer state via IVON_INIT_METHOD=trained in the companion code's run_rl.sh.
Citation
@misc{venkatkrishna2026parameterexploration,
title={Parameter Exploration for RLVR via Variational Learning},
author={Vatsal Venkatkrishna and Nico Daheim and Iryna Gurevych},
year={2026},
eprint={2608.09805},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2608.09805},
}