CoolFace
Modelpublic

BayesRL/Llama3.1-IVON-SFT-8B

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes7.3kdownloads
Model Card

Qwen2.5Math-IVON-SFT-7B

📄 Paper: Parameter Exploration for RLVR via Variational Learning · arXiv:2608.09805

📦 Code: insait-institute/c3po

Qwen2.5-Math 7B supervised-fine-tuned with the variational optimizer IVON, from the paper "Parameter Exploration for RLVR via Variational Learning".

This is a warm-start checkpoint: SFT'ing with IVON yields not just point weights but an approximate Gaussian posterior over them (a mean and a diagonal Hessian/precision estimate). That posterior is the learned prior used to seed the 3PO RLVR runs (B3PO / M3PO / C3PO), where weight perturbations sampled from it drive parameter-space exploration.

Training

Foundation modelQwen/Qwen2.5-Math-7B
StageWarm-start SFT
DataLlama-Nemotron Post-Training Dataset (SFT subset)
OptimizerIVON, lr 50.0, ESS (λ) 1e10
Hardware8× NVIDIA H200 (144 GB)

Usage

Loads as a standard causal LM:

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("BayesRL/Qwen2.5Math-IVON-SFT-7B")
tok = AutoTokenizer.from_pretrained("BayesRL/Qwen2.5Math-IVON-SFT-7B")

To use it as the warm-start prior for 3PO RLVR, load the IVON optimizer state via IVON_INIT_METHOD=trained in the companion code's run_rl.sh.

Citation

bibtex
@misc{venkatkrishna2026parameterexploration,
      title={Parameter Exploration for RLVR via Variational Learning}, 
      author={Vatsal Venkatkrishna and Nico Daheim and Iryna Gurevych},
      year={2026},
      eprint={2608.09805},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2608.09805}, 
}