CoolFace
Modelpublic

NebulaPixel/SummOrchestra-Qwen3-8B-GRPO-BRLP-SAMSUM

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
1likes36downloads
Model Card

dadastory/SummOrchestra-Qwen3-8B-GRPO-BRLP-SAMSUM is a dialogue-summarization model trained on the SamSum dataset through a two-stage pipeline combining Supervised Fine-Tuning (SFT) and GRPO + BRLP preference optimization. This design yields a model with higher faithfulness, reduced hallucination, and stronger alignment with human preference judgments.

Performance Comparison

The model consistently outperforms standard SamSum-style summarizers in both loyalty and human preference metrics.

<p align="center"> <img src="./barplot_comparison.png" alt="Barplot Comparison" width="70%"> </p>

<p align="center"> <img src="./kdefaithfulnesscomparison.png" alt="Kde Faithfulness Comparison" width="70%"> </p>

<p align="center"> <img src="./kdehumanpreference_comparison.png" alt="Kde Human Preference Comparison" width="70%"> </p>

Training Overview

1. Supervised Fine-Tuning (SFT)

  • —Trained on SamSum to establish strong grounding and structural summarization ability.

2. GRPO + BRLP Preference Optimization

  • —GRPO introduces preference-driven refinement.
  • —BRLP (Behavioral Reward Learning) provides stable reward shaping.

Evaluation & Alignment

Loyalty / Faithfulness Evaluation

Faithfulness is evaluated using vectara/hallucination_evaluation_model, ensuring:

  • —Lower hallucination rate
  • —Higher semantic grounding
  • —More accurate retention of source dialogue meaning

Human Preference Evaluation

Human-likeness and preference alignment are assessed using Qwen/WorldPM-72B-RLHFLow, focusing on:

  • —Natural summarization style
  • —Alignment with human judgment patterns
  • —Improved preference scoring relative to SamSum baselines

Key Characteristics

  • —Base Model: Qwen3-8B
  • —Dataset: SamSum dialogues
  • —Training Stages: SFT + GRPO + BRLP
  • —Goals:
  • —Enhance faithfulness
  • —Improve key information coverage
  • —Reduce hallucination
  • —Increase human preference alignment