promotion/Llama-3.1-8B-SafeRLHF-NBPO-600updates
0346
NBPO on SafeRLHF, 600 updates
Budget ablation on SafeRLHF: both the held-out fit and the win rate prefer the 300-update arm (nMSE 0.9762 vs 0.9591; worst 0.688 vs 0.696).
- Panel: SafeRLHF
- Objectives: helpfulness, harmlessness
- Backbone / reference policy:
meta-llama/Llama-3.1-8B-Instruct - Training budget: 300 optimiser updates, global batch 16
- Reported in: Nash Bargaining Preference Optimization (NBPO), Appendix: optimiser-budget sensitivity
Evaluation protocol: independent objective-wise win rate against the common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, both presentation orders.
