CoolFace
Modelpublic

promotion/Llama-3.1-8B-SafeRLHF-MaxMinRLHF-baseline

sourceHugging Facellama3.1updated 22d agoView on Hugging Face
0likes291downloads
Model Card

MaxMin-RLHF on SafeRLHF

MaxMin-RLHF (Chakraborty et al., ICML 2024), Algorithm 1 with the panel's objectives as the known groups: three alternating rounds of 100 updates, each on the objective with the lowest reference-standardised utility.

  • —Panel: SafeRLHF
  • —Objectives: helpfulness, harmlessness
  • —Backbone / reference policy: meta-llama/Llama-3.1-8B-Instruct
  • —Training budget: 300 optimiser updates, global batch 16
  • —Reported in: Nash Bargaining Preference Optimization (NBPO), Table 2 (primary cross-method evaluation)

Evaluation protocol: independent objective-wise win rate against the common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, both presentation orders.