CoolFace
Modelpublic

just1nseo/llama31-8b-if-rlvr-anchor-pyx01

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
Model Card

just1nseo/llama31-8b-if-rlvr-anchor-pyx01

GRPO / IF-RLVR checkpoints for meta-llama/Llama-3.1-8B-Instruct, trained with the bidirectional anchor reward (p(y|x) anchor coefficient 0.1).

Each checkpoint is a complete HF model directory in its own subfolder:

python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("just1nseo/llama31-8b-if-rlvr-anchor-pyx01", subfolder="global_step_91")
tok = AutoTokenizer.from_pretrained("just1nseo/llama31-8b-if-rlvr-anchor-pyx01", subfolder="global_step_91")

Checkpoints

subfolderepochuploaded (UTC)
global_step_9112026-08-04 21:52:20
global_step_18222026-08-05 06:32:30
global_step_27332026-08-05 15:07:39
global_step_36442026-08-05 23:26:48
global_step_45552026-08-07 12:10:35
global_step_54662026-08-07 20:17:49

Training configuration

experimentllama31_8b_grpo_nonthink_pyx01_t8banchor_s8b_b1024_c1
anchor cacheif_ref_anchor_teacher_llama31_8b_instruct_nonreason_train_seed1_val512_scored_by_llama31_8b_instruct_topp095_topk20.json
steps / epoch91
train batch size1024
rollout n8
max prompt / response2048 / 2048
wandbifif/verl_if_rlvr