stindardlogic/summarization-sft-100k
Summarization SFT (100K) 100,000 ShareGPT-format conversations covering document summarization across 9 source types and 5 summary styles. Trains models to summarize professional documents the way an expert human analyst would — identifying what matters, choosing the right format, and calibrating length to the task. Motivation Summarization is one of the most commercially deployed LLM capabilities, yet most summarization datasets train on news articles only. Real… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/summarization-sft-100k.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face