promotion/Llama-3.1-8B-TLDR-RewardedSoups-baseline
0323
Rewarded Soups on TL;DR
Rewarded Soups (Rame et al., NeurIPS 2023): one expert per objective from the shared initialisation, linearly interpolated; the interpolation weight was chosen on validation prompts only.
- Panel: TL;DR
- Objectives: coverage, faithfulness, conciseness, helpfulness
- Backbone / reference policy:
meta-llama/Llama-3.1-8B-Instruct - Training budget: 300 optimiser updates, global batch 16
- Reported in: Nash Bargaining Preference Optimization (NBPO), Table 2 (primary cross-method evaluation)
Evaluation protocol: independent objective-wise win rate against the common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, both presentation orders.
