CoolFace
Modelpublic

promotion/Llama-3.1-8B-SafeRLHF-NBPO-stage2

sourceHugging Facellama3.1updated 20d agoView on Hugging Face
0likes334downloads
Model Card

NBPO on SafeRLHF after a second outer stage

pi2 from a genuine second stage of Algorithm 1: the pool regenerated from pi1, the training matrix rejudged in full (44,000 cells), the dual re-solved and the policy refit for 300 updates. Accepted by the Algorithm 1 gate (held-out surpluses 0.240 / 0.148) but it does not improve on pi_1: +0.003 on the gated surplus and -0.008 on the reported worst-objective win rate.

  • —Panel: SafeRLHF
  • —Objectives: helpfulness, harmlessness
  • —Backbone / reference policy: meta-llama/Llama-3.1-8B-Instruct
  • —Training budget: 300 optimiser updates, global batch 16
  • —Reported in: Nash Bargaining Preference Optimization (NBPO), Appendix: A Second Outer Stage

Evaluation protocol: independent objective-wise win rate against the common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, both presentation orders.