CoolFace
Modelpublic

promotion/Llama-3.1-8B-TLDR-HTMNPO-helpfulness

sourceHugging Facellama3.1updated 26d agoView on Hugging Face
0likes18downloads
Model Card

Llama-3.1-8B-TLDR-HTMNPO-helpfulness

Single-objective corner on the TL;DR panel: all weight on helpfulness.

Trained from meta-llama/Llama-3.1-8B-Instruct, which is also the reference policy and the initialisation. Objectives are scored by a prompted Qwen3-32B preference oracle, each pair queried in both presentation orders and swap-averaged. Within a panel every arm shares one response pool, one optimizer and a 300-step budget, and differs only in how the objectives are aggregated, so a difference between two arms is attributable to the aggregation rule.

Held-out surplus over the reference on 100 prompts, population scale \(Ak = Pk - 1/2\):

objectivesurplus
coverage+0.0733
faithfulness-0.0498
conciseness-0.0697
helpfulness+0.0964
minimum-0.0697
average+0.0125

Bootstrap intervals and paired significance tests for this panel are in the paper's appendix. Benchmark generations for the UltraFeedback arms are at `promotion/nbpo-benchmark-generations`.

Built with Llama. Use is subject to the Llama 3.1 Community License.