sangyon/llama31_tulu3_8b_dpo_grpo_nonthink_intentcheck_tulu3_8b_dpo
0288
Tulu 3 8B DPO — IntentCheck GRPO
Hugging Face-format export of the global_step_364 checkpoint from the non-reasoning Tulu IntentCheck experiment.
- Initial policy:
allenai/Llama-3.1-Tulu-3-8B-DPO. - IntentCheck judge: the fixed Tulu 3 8B DPO model.
- Reward: programmatic constraint reward, plus
0.1when an eligible response passes IntentCheck. Non-passing responses do not receive the additional bonus. - Training: GRPO, four epochs, batch size1024, eight rollouts per prompt, maximum response length2048 tokens.
- Actor KL coefficient:
0.001. - Export: four FSDP model shards merged into BF16 safetensors. Optimizer state is not included.
- Validation: tensor finiteness, local model/tokenizer loading, and a finite CPU forward pass were checked. No new benchmark evaluation is claimed for this export.
Experiment: llama31_tulu3_8b_dpo_grpo_nonthink_intentcheck_tulu3_8b_dpo_bonus01_b1024_c1_t1_2k_e4_s91_20260907b.
Final resumed training run: https://wandb.ai/ifif/verlifrlvr/runs/ti4r0908b
The model is subject to the Llama3.1 Community License and applicable acceptable-use policy inherited from its base model.
