CoolFace
Modelpublic

just1nseo/llama31-tulu3-8b-dpo-if-rlvr-anchor-judge

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
Model Card

just1nseo/llama31-tulu3-8b-dpo-if-rlvr-anchor-judge

GRPO / IF-RLVR checkpoints for meta-llama/Llama-3.1-8B-Instruct, trained with the bidirectional anchor reward (p(y|x) anchor coefficient 0.1 + gpt-oss-120b verifier fallback bonus 0.1 (threshold 5)).

Each checkpoint is a complete HF model directory in its own subfolder:

python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("just1nseo/llama31-tulu3-8b-dpo-if-rlvr-anchor-judge", subfolder="global_step_91")
tok = AutoTokenizer.from_pretrained("just1nseo/llama31-tulu3-8b-dpo-if-rlvr-anchor-judge", subfolder="global_step_91")

Checkpoints

subfolderepochuploaded (UTC)
global_step_9112026-08-10 02:10:21
global_step_18222026-08-10 16:16:45
global_step_27332026-08-11 06:40:08
global_step_36442026-08-11 20:49:32

Training configuration

experimentllama31_tulu3_8b_dpo_grpo_nonthink_anchor_pyx01_llmverifier_gptoss120b_fallback_bonus01_threshold5_b1024_c1_t1_2k
anchor cacheif_ref_anchor_teacher_llama31_tulu3_8b_dpo_nonreason_train_seed1_val512_t1_p095_r2048_scored_by_llama31_tulu3_8b_dpo.json (v3, 93993 rows all complete)
steps / epoch91
train batch size1024
rollout n8
max prompt / response2048 / 2048
wandbifif/verl_if_rlvr