promotion/Llama-3.1-8B-SafeRLHF-NBPO-eta0.1
0282
NBPO on SafeRLHF at the selected eta_t
The SafeRLHF NBPO arm retrained at the etat the pooled-entropy rule selects (0.1) instead of the released etat = 1, reusing the same precomputed pair artifact; only eta_t differs.
- Panel: SafeRLHF
- Objectives: helpfulness, harmlessness
- Backbone / reference policy:
meta-llama/Llama-3.1-8B-Instruct - Training budget: 300 optimiser updates, global batch 16
- Reported in: Nash Bargaining Preference Optimization (NBPO), Table 2 (primary cross-method evaluation)
Evaluation protocol: independent objective-wise win rate against the common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, both presentation orders.
