NebulaPixel/SummOrchestra-Qwen3-8B-GRPO-BRLP-SAMSUM
dadastory/SummOrchestra-Qwen3-8B-GRPO-BRLP-SAMSUM is a dialogue-summarization model trained on the SamSum dataset through a two-stage pipeline combining Supervised Fine-Tuning (SFT) and GRPO + BRLP preference optimization. This design yields a model with higher faithfulness, reduced hallucination, and stronger alignment with human preference judgments.
Performance Comparison
The model consistently outperforms standard SamSum-style summarizers in both loyalty and human preference metrics.
<p align="center"> <img src="./barplot_comparison.png" alt="Barplot Comparison" width="70%"> </p>
<p align="center"> <img src="./kdefaithfulnesscomparison.png" alt="Kde Faithfulness Comparison" width="70%"> </p>
<p align="center"> <img src="./kdehumanpreference_comparison.png" alt="Kde Human Preference Comparison" width="70%"> </p>
Training Overview
1. Supervised Fine-Tuning (SFT)
- Trained on SamSum to establish strong grounding and structural summarization ability.
2. GRPO + BRLP Preference Optimization
- GRPO introduces preference-driven refinement.
- BRLP (Behavioral Reward Learning) provides stable reward shaping.
Evaluation & Alignment
Loyalty / Faithfulness Evaluation
Faithfulness is evaluated using vectara/hallucination_evaluation_model, ensuring:
- Lower hallucination rate
- Higher semantic grounding
- More accurate retention of source dialogue meaning
Human Preference Evaluation
Human-likeness and preference alignment are assessed using Qwen/WorldPM-72B-RLHFLow, focusing on:
- Natural summarization style
- Alignment with human judgment patterns
- Improved preference scoring relative to SamSum baselines
Key Characteristics
- Base Model: Qwen3-8B
- Dataset: SamSum dialogues
- Training Stages: SFT + GRPO + BRLP
- Goals:
- Enhance faithfulness
- Improve key information coverage
- Reduce hallucination
- Increase human preference alignment
