just1nseo/llama31-8b-if-rlvr-anchor-pyx01
0
just1nseo/llama31-8b-if-rlvr-anchor-pyx01
GRPO / IF-RLVR checkpoints for meta-llama/Llama-3.1-8B-Instruct, trained with the bidirectional anchor reward (p(y|x) anchor coefficient 0.1).
Each checkpoint is a complete HF model directory in its own subfolder:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("just1nseo/llama31-8b-if-rlvr-anchor-pyx01", subfolder="global_step_91")
tok = AutoTokenizer.from_pretrained("just1nseo/llama31-8b-if-rlvr-anchor-pyx01", subfolder="global_step_91")