CoolFace
Modelpublic

sangyon/llama31_tulu3_8b_dpo_grpo_nonthink_intentcheck_tulu3_8b_dpo

sourceHugging Facellama3.1updated 19d agoView on Hugging Face
0likes288downloads
Model Card

Tulu 3 8B DPO — IntentCheck GRPO

Hugging Face-format export of the global_step_364 checkpoint from the non-reasoning Tulu IntentCheck experiment.

  • —Initial policy: allenai/Llama-3.1-Tulu-3-8B-DPO.
  • —IntentCheck judge: the fixed Tulu 3 8B DPO model.
  • —Reward: programmatic constraint reward, plus 0.1 when an eligible response passes IntentCheck. Non-passing responses do not receive the additional bonus.
  • —Training: GRPO, four epochs, batch size1024, eight rollouts per prompt, maximum response length2048 tokens.
  • —Actor KL coefficient: 0.001.
  • —Export: four FSDP model shards merged into BF16 safetensors. Optimizer state is not included.
  • —Validation: tensor finiteness, local model/tokenizer loading, and a finite CPU forward pass were checked. No new benchmark evaluation is claimed for this export.

Experiment: llama31_tulu3_8b_dpo_grpo_nonthink_intentcheck_tulu3_8b_dpo_bonus01_b1024_c1_t1_2k_e4_s91_20260907b.

Final resumed training run: https://wandb.ai/ifif/verlifrlvr/runs/ti4r0908b

The model is subject to the Llama3.1 Community License and applicable acceptable-use policy inherited from its base model.