just1nseo/olmo3-7b-think-if-rlvr-anchor-pyx01-pp0-8k
0
OLMo 3 7B Think — IF-RLVR anchor checkpoints
BF16 inference checkpoints for allenai/Olmo-3-7B-Think-DPO, trained with GRPO on instruction-following constraints and a p(y|x) anchor reward coefficient of 0.1.
Each checkpoint is a complete Transformers model in its own subfolder:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "just1nseo/olmo3-7b-think-if-rlvr-anchor-pyx01-pp0-8k"
subfolder = "global_step_364"
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=subfolder)
model = AutoModelForCausalLM.from_pretrained(
repo,
subfolder=subfolder,
torch_dtype=torch.bfloat16,
device_map="auto",
)OLMo 3 requires a Transformers version with native olmo3 support (the training runtime used Transformers 4.57.1).
Checkpoints
Training configuration
The full FSDP trainer checkpoints, including optimizer and data-loader state, are retained separately. This repository contains merged BF16 model weights intended for inference and evaluation.
