datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-summary-dataset
دیتاست خلاصهسازی فارسی
این دیتاست شامل ۱۰۶۳ نمونه برای خلاصهسازی معنایی متون فارسی است.
فرمت داده
هر نمونه شامل سه فیلد است:
instruction: دستور خلاصهسازی (بیش از ۴۰ دستور مختلف)
input: متن ورودی (بین ۱۰۰ تا ۴۰۰ کلمه)
output: خلاصه تولید شده (بین ۳۰ تا ۱۰۰ کلمه)
تعداد نمونهها
۱۰۶۳ نمونه با کیفیت بالا
حوزههای پوشش داده شده
اقتصادی (بانک، بورس، بیمه، تولید)
سیاسی (داخلی، خارجی، مذاکرات)
اجتماعی (آسیبشناسی، رفاه، اشتغال)
علمی و… See the full description on the dataset page: https://huggingface.co/datasets/msdos34/persian-summary-dataset.Book_Summary_Chinese
中文图书总结数据集
每个样本包含:
图书的一个章节、此章节的总结、图书名字,可以训练模型总结长文本的能力。数据主要来自较为著名的中文版小说。
story-summary
Short Story Summarization Dataset
Short stories from agentlans/euclaise-WritingPromptsX
Summarized using Qwen/Qwen3.5-9B with the following prompt:
Summarize the short story below in a single well-written paragraph that captures the main plot, key characters, central conflict, and resolution while preserving the original meaning and tone. Focus only on the most important details, avoid unnecessary specifics or minor subplots, and do not add interpretations or information that… See the full description on the dataset page: https://huggingface.co/datasets/k007-a/story-summary.story-summary
Short Story Summarization Dataset
Short stories from agentlans/euclaise-WritingPromptsX
Summarized using Qwen/Qwen3.5-9B with the following prompt:
Summarize the short story below in a single well-written paragraph that captures the main plot, key characters, central conflict, and resolution while preserving the original meaning and tone. Focus only on the most important details, avoid unnecessary specifics or minor subplots, and do not add interpretations or information that… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/story-summary.summary-keywordshigh-quality-summary
Data from agentlans/high-quality-text sample_k10000 configuration
Summaries generated using google/gemma-3-12b-it
Summaries rewritten using agentlans/granite-3.3-2b-refiner
Rewritten summaries checked against the original text using ibm-granite/granite-3.3-8b-instruct
turkish-doc-summary-review-analysis-30k-jsonl
⚠️ Superseded by v2
Bu v1 dataset'te exact duplicate yoktu; ancak belge ve cevap şablonları fazla tekrar ediyordu.
Güncel v2 sürümünü kullanın:
https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl-v2
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code:… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl.turkish-doc-summary-review-analysis-30k-jsonl-v2
Turkish Document Summary Review Analysis 30K JSONL v2
Tek dosya: train.jsonl.
Bu v2 sürümü, v1'de görülen tekrar sorununu çözmek için yeniden üretildi:
Daha fazla belge türü: proje önerisi, tutanak, denetim notu, şikâyet dosyası, politika taslağı, saha raporu, bütçe değerlendirmesi, risk kayıt formu, karar destek belgesi, olay inceleme raporu vb.
Daha fazla alt görev: 30 farklı task_type.
Exact duplicate + semantic template duplicate kontrolü.
Cevap şablonları belgeye özel risk… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl-v2.subtitle-summary-dataset
Subtitle Summary & Keyword Dataset (YouTube)
국내 방송사 유튜브 자막 기반 스트리밍(누적) 요약 + 유튜브 검색 키워드 데이터셋.
방송을 5분 단위로 진행하며, 매 5분마다 (직전 누적 요약 + 새 5분 자막)으로 방송 대표 한 문장 요약을 갱신하고 검색 키워드를 생성합니다.
교사 모델 Qwen3-235B-A22B-Instruct-2507-FP8 (vLLM, 자막-only). HBKenerzai/LGUplus_summary_keyword와 동일 스키마.
규모
split
회차
레코드(5분)
train
2,641
10,159
validation
294
971
합계
2,935
11,130
약 563시간 · 110개 채널 · 요약 중앙값 58자 · 도메인(유튜브 카테고리): Entertainment 2982 · Travel & Events 2081 · News &… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-dataset.Instruct-Summary
Dataset Card for "Instruct-Summary"
This dataset is a combination of kmfoda/booksum, samsum, mosaicml/dolly_hhrlhf and yahma/alpaca-cleaned.
tw-law-context-summary
Dataset Card for tw-law-context-summary
本資料集為中華民國(臺灣)法規條文之 LLM 摘要集,每筆樣本提供「法規名稱」對應的條列式摘要文字,可作為法規 RAG 系統的索引/簡介,或法律 chatbot 的初步說明資料。
Dataset Details
Dataset Description
資料以法規為單位,由 LLM 對每部法規生成「立法依據、規範重點、施行日期、適用範圍、廢止狀態」等結構化摘要。每筆樣本欄位:
text:摘要內容(條列式)。
name:法規名稱。
abandon_note:廢止/修訂註記,若為現行法規則為空字串。
token_count / word_count:保留為字串欄位(部分樣本為空)。
主要使用者為法規檢索、法律入門教育場景。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional Chinese
License: cc-by-nc-sa-4.0… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-law-context-summary.tw-processed-related-law-article-summary
Dataset Card for tw-processed-related-law-article-summary
本資料集為中華民國(臺灣)法規「條文之間引用關係」的 LLM 摘要:每筆樣本針對某條條文與其引用之相關條文,提供結構化的關係分析(直接引用、技術延伸、規範一致性、實務操作、廢止情況等)。
Dataset Details
Dataset Description
tw-processed-related-law-article 已將每條條文與其引用之條文串接在一起;本資料集則進一步以 LLM 解析「該條與所引用條之間的關係」並輸出條列式摘要,內容通常包含:
直接引用與依據關係
技術規範的延伸
法規一致性
實務操作影響
廢止情況
可作為法規 RAG 系統的「條文關係層」索引,協助下游模型理解條文之間的依存。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional Chinese
License: cc-by-nc-sa-4.0… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-related-law-article-summary.Haruhi-Dialogue-Speaker-Extract-And-Summary之前的 silk-road/Haruhi-Dialogue-Speaker-Extract 要求模型输出json格式,并且采取了CoT策略,感觉有一些难了
这一次把总结和抽取拆分成了两个任务
并且抽取的格式改为了csv格式。
subtitle-summary-lgu
Broadcast Subtitle Summary & Keyword — (LGUplus source, our pipeline)
HBKenerzai/LGUplus_summary_keyword의 방송 자막을 우리 형식으로 재구성한 뒤, **우리 파이프라인
(Qwen3-235B-A22B-Instruct-2507-FP8, vLLM, 자막-only)**으로 5분 단위 누적 요약 + 검색 키워드를
다시 생성한 데이터셋입니다. 원천 자막·메타 저작권은 원 출처(AI Hub / 방송사)에 있습니다.
규모
split
회차
레코드(5분)
train
16,609
118,608
validation
1,846
13,177
합계
18,455
131,785
약 9,025시간 분량 · 요약 길이 중앙값 63자 · 도메인: 생활정보 9528 · 정보/토크 9458 · 퀴즈/게임 8123 · 푸드/요리 7221 · 토론/대담… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-lgu.subtitle-summary-sft
Subtitle Summary & Keyword — SFT (YouTube)
jungsanghyun/subtitle-summary-dataset를 SFT 학습용 messages 형식으로 가공.
형식 (system 없음, 단일 턴)
입력(user): [이전 요약] {last_summary}\n[자막] {5min_script}
출력(assistant): [요약] {한 문장}\n[키워드] {검색어} — 라벨 2줄(소형 모델 친화).
split
examples
train
10,159
validation
971
chat_template.jinja({% generation %})로 assistant만 loss. 파싱 \[요약\]\s*(.+) / \[키워드\]\s*(.+).
라이선스
cc-by-nc-4.0, 연구·비상업.
subtitle-summary-testset
Subtitle Summary & Keyword — Test set (YouTube)
방송 유튜브 자막 기반 누적 요약 + 검색 키워드 태스크의 평가용 테스트셋입니다.
jungsanghyun/subtitle-summary-dataset의 train/validation과 겹치지 않는 별도 방송으로, 동일 파이프라인(Qwen3-235B, vLLM, 자막-only)으로 생성했습니다.
규모
split
회차
레코드(5분)
test
114
519
약 25시간 · 요약 중앙값 59자 · 도메인: Entertainment 236 · News & Politics 103 · Pets & Animals 87 · Travel & Events 34 · Education 31 · Music 28
스키마
program_name, last_summary, 5min_script(입력) →… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-testset.LGUplus_summary_keyword
LGUplus Broadcast Streaming Summary & Keyword Dataset
한국어 방송 자막 기반 스트리밍(누적) 요약 + 키워드 추출 데이터셋입니다.
방송을 5분 단위로 진행하며, 매 5분마다 (직전까지의 요약 + 새 5분 자막)을 받아
방송 전체를 대표하는 한 문장 요약으로 갱신하고, 해당 5분의 핵심 키워드 1개를 뽑는 태스크를 위해 구축되었습니다.
데이터 구조
각 레코드 = 한 방송의 한 5분 구간(step).
필드
설명
program_name
프로그램명
last_summary
직전 시점까지의 누적 요약 (첫 청크는 "") — 입력
5min_script
이번 5분 구간 자막 ([역할] 발화 형식) — 입력
current_summary
갱신된 방송 전체 대표 요약 (한 문장) — 출력
current_5min_keyword
이번 5분 구간 핵심 키워드 (1개) — 출력… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/LGUplus_summary_keyword.subtitle-summary-lgu-sft
Broadcast Subtitle Summary & Keyword — SFT (LGUplus source)
jungsanghyun/subtitle-summary-lgu를 SFT 학습용 messages 형식으로 가공한 버전입니다.
형식 (system 없음, 단일 턴)
입력(user): [이전 요약] {last_summary}\n[자막] {5min_script}
출력(assistant): [요약] {한 문장}\n[키워드] {검색어} — 라벨 2줄(소형 모델 친화·파싱 용이).
split
examples
train
118,608
validation
13,177
generation만 학습
chat_template.jinja({% generation %} 블록)로 assistant 응답만 loss. TRL SFTConfig(assistant_only_loss=True… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-lgu-sft.subtitle-summary-testset-sft
Subtitle Summary & Keyword — Test set SFT (YouTube)
jungsanghyun/subtitle-summary-testset를 messages 형식으로 가공한 평가용 버전.
형식 (system 없음, 단일 턴)
입력(user): [이전 요약] {last_summary}\n[자막] {5min_script}
출력(assistant): [요약] {한 문장}\n[키워드] {검색어}
split
examples
test
519
파싱 \[요약\]\s*(.+) / \[키워드\]\s*(.+). chat_template.jinja 동봉.
라이선스
cc-by-nc-4.0, 연구·비상업.
high-quality-summary-v2
High Quality Long Text Summarization Dataset
Input texts from agentlans/high-quality-text-long sample_k10000 config
Summaries generated by google/gemma-3-12b-it
Summaries rewritten by agentlans/granite-3.3-2b-reviser
