DISLab/SummLlama3-70B
<div align="center"> <b style="font-size: 40px;">SummLlama3-70B</b> </div>
Are you looking for a summarizer that can generate more human-preferred summaries across multiple domains?
Our SummLlama3-70B could be exactly what you need!
SummLlama3-70B is initialized from Llama3-70B-Instruct, with additional training using Direct Preference Optimization (DPO) based on large-scale (over 100K) summarization feedback.
The feedback encompasses a wide range of input documents, from short to lengthy texts, including both dialogue and non-dialogue formats, and spans across seven distinct domains:
- Four non-dialouge domains: News, Lifestyle, Report, Medical
- Three dialogue domains: Daily Life, Interview, Meeting
This is automated evaluation results:
Please refer to our paper to catch up how to exploit LLM-generated feedback in the context of text summarization.
SummLlama3-70B,
https://huggingface.co/DISLab/SummLlama3-8B
https://huggingface.co/DISLab/SummLlama3-70B
SummLlama3.1-Series
https://huggingface.co/DISLab/SummLlama3.1-8B
https://huggingface.co/DISLab/SummLlama3.1-70B
SummLlama3.2-Series
https://huggingface.co/DISLab/SummLlama3.2-3B
Recommended Prompt for Text Summarization:
We recommend to use the prompt below to get the summary, since we trained the model using this.
def format_chat_template(document):
instruction = "Please summarize the input documnet."
row_json = [{"role": "user", "content": f"Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\n{instruction}\n\n### Input:\n{document}\n\n### Response:\n"}]
return tokenizer.apply_chat_template(row_json, tokenize=False)Here is a brief overview of our summarizer:
Rather than relying on expensive human feedback, we utilize high-quality, multi-dimensional, and fine-grained feedback generated by large language models (LLMs).
This model excels at faithfulness, completeness, and conciseness, which are the three human-preferred aspects to judge what is a good summarizer.
- Faithfulness: a summarizer does not manipulate the information in the input text and add any information not directly inferable from the input text.
- Completeness: a summarizer ensures the inclusion of all key information from the input text in the output summary.
- Conciseness: a summarizer refrains from incorporating information outside the key information in the output, maintaining a succinct and focused summary.
Based on our comprehensive evaluation, which included both human and automated assessments of summary quality, SummLlama3 demonstrated significant improvements over the original Llama3 series.
Here is the results:
Human Evaluation
Autoamted Evaluation using FineSurE
Example
See an example how the summary improved by SummLlama3-8B over Llama3-8/70B-Instruct on the document below:
The summary of SummLlama3-8B can be considered a much human-preferred summary for the following reasons:
Core Focus: The summary accurately captures the main theme of the conversation, which revolves around the Thanksgiving dinner arrangements. It highlights how the two people confirm plans, discuss what to bring, and finalize the decision for Person2 to bring wine instead of pie. This maintains the core context.
Inclusion of Key-facts: The summary covers the important details of the conversation, including Person2's initial offer to bring dessert (pumpkin pie) and the shift to bringing wine due to another family member handling dessert. Other summaries tend to overlook or simplify this progression, while SummLlama3-8B fully captures the interaction’s key events.
Clarity and Conciseness: The summary is structured in a straightforward, concise manner, effectively summarizing the conversation without unnecessary details. It presents the flow and outcome of the discussion clearly, making it easy for readers to understand. The logical order of events is maintained, ensuring a smooth narrative.
Accurate Role Depiction: The summary clearly identifies Person1 as the host and Paul (Person2) as the guest, which helps clarify their relationship and the nature of the conversation. This distinction is more explicit in SummLlama3-8B compared to other summaries, which might leave these roles more ambiguous.
