datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
summarize_from_feedback_tldr_3_filteredThis is the query dataset taken directly from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
light-batch-summarize-dialogue
Light dataset
Dialogues are preprocessed into a form:
<Character name>: <character line>
...
<Character name>: <character line>
Summarize the document
clinical-summarizer-sft
Clinical Note Summarizer (SOAP Format)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Long clinical notes → structured SOAP summaries
Why download this
Automate clinical documentation. Reduce physician burnout by summarizing visit notes into Subjective / Objective / Assessment / Plan format.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/clinical-summarizer-sft.openai-summarize-tldr
Summarize TL;DR Filtered Dataset
This is the version of the dataset used in https://arxiv.org/abs/2009.01325.
If starting a new project we would recommend using https://huggingface.co/datasets/openai/summarize_from_feedback.
For more information see https://github.com/openai/summarize-from-feedback and for the original TL;DR dataset see https://huggingface.co/datasets/webis/tldr-17.
Korean-1930-Novel-Scene-Summarize
한국 저작권 만료 소설에 대한 씬 분리 및 요약 데이터 셋
원천 데이터 출처: https://gongu.copyright.or.kr/gongu/wrt/wrtCl/listWrtText.do?menuNo=200019
총 96개 소설 수집 및 전처리
한자가 많은 소설 제외
한자 제거, 띄어쓰기 전처리 수행
씬 분리
사용 모델: Gemini-1.5-Flash
(띄어쓰기 포함) 100자 이상, 1200자 미만으로 적절한 문장에서 씬 단위로 분리하도록 지시
총 12,108씬 생성
요약
사용 모델: Gemini-1.5-Flash(때때로 GPT-4o)
각 Scene에서 인물, 주요 소품, 사건을 추출하고, 요약(scenario)을 생성하도록 함
summarize-from-feedback-prosummarize_from_feedback.kr영문 데이터셋 summarize_from_feedback을 한영 번역 모델인 Gugugo-koen를 이용하여 번역함.
원본은 batch3.json에서 batch22.json까지 있지만, 시간 관계상 batch9.json까지만 작업하고 중지함. 모든 데이터가 필요한 분은 원본을 참고해서 후속 작업 요망.
summarize_from_feedback : https://huggingface.co/datasets/openai/summarize_from_feedback
Gugugo-koen : https://huggingface.co/squarelike/Gugugo-koen-7B-V1.1-AWQ
smol_summarize_pt
SMOL Summarize PT
This dataset is the translated version of the smol-summarize subset of the HuggingFaceTB/smoltalk.
Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
Note: This dataset comprises machine translated content and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol_summarize_pt.Youtube-transcript-Summarizeropenai_summarize_comparisons_relabelThe dataset is a relabel dataset of the CarperAI/openai_summarize_comparisons dataset.
The annotators are Reward Model trained on the train split of CarperAI/openai_summarize_comparisons dataset based on the llama-2-7b .
The annotation python script is as follows:
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
from datasets import load_dataset
from scipy.special import expit
import numpy as np
import random
import json
from tqdm import tqdm
import… See the full description on the dataset page: https://huggingface.co/datasets/chadlzx/openai_summarize_comparisons_relabel.summarize-long-texts
Summarize long texts
This dataset provides long prompts intended to use for testing language models with long inputs with different sizes.
Columns
id: numerical id
source: the source of the prompt text
length: indication of length of the prompt
text: the text of the prompt
Item lengths
Exact prompt length depend on the tokenizer for the model, and can differ quite a bit with different tokenizers.
The lengths included in the dataset should be seen as… See the full description on the dataset page: https://huggingface.co/datasets/helenai/summarize-long-texts.health_summarizetldrsumthink_responses_summarizedSee: https://huggingface.co/datasets/G-reen/sumthink
This contains the summarized GLM 4.7 nvfp4 thinking traces as well as the source data.
AP-News-2024-CGPT-Summarize-ShareGPTThe AP News dataset, run through ChatGPT (gpt-3.5-turbo) to get summaries.
All use the same system prompt; "You summarize text. Ensure that your summaries effectively capture key points, while being concise."
Currently not all of the articles from the dataset are summarized, since I keep hitting "You've reached our limit of messages per hour. Please try again later."
chansung_new_summarize_synth_ds-gpt4o-1371c6b-ShareGPTSummarized-Attractionszoho-1b-summarize-kimi-k2-thinking-65k-w-reasoningeron-mails-summarizeda
summarize-newsAP-News-CGPT-Summarize-Not-Cucked-PreferenceShareGPTPubmedQA_5_WITH_RELATION_SUMMARIZE_vsimilarity_bothsummarize_llama3.2Paper_1_BioASQBlurb_5_formated_metadata_MetadataMethod.WITH_RELATION_SUMMARIZE_vqc_both-datasethotpot_qa_summarizedBioASQBlurb_5_WITH_RELATION_SUMMARIZE_vqc_bothBioASQBlurb_5_WITH_RELATION_SUMMARIZE_vsimilarity_bothBioASQ_5_WITH_RELATION_SUMMARIZE_vqcBioASQ_5_WITH_RELATION_SUMMARIZE_vsimilarityPubmedQA_5_WITH_RELATION_SUMMARIZE_vqc_both
