promotion/Qwen3-8B-Panacea-baseline
0281
Panacea on the Qwen3-8B backbone
Panacea (Zhong et al., NeurIPS 2024): SVD-LoRA with k = 8 preference-agnostic singular values and the preference vector injected as the remaining ones, LS aggregation, exported at lambda = (0.15, 0.55, 0.15, 0.15), chosen on 100 validation prompts only.
- Panel: UltraFeedback / Qwen3-8B
- Objectives: instruction following, truthfulness, honesty, helpfulness
- Backbone / reference policy:
Qwen/Qwen3-8B - Training budget: 300 optimiser updates, global batch 16
- Reported in: Nash Bargaining Preference Optimization (NBPO), Table 1 (general-capability evaluation)
Evaluation protocol: independent objective-wise win rate against the common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, both presentation orders.
