CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01msdos34 /persian-summary-dataset دیتاست خلاصه‌سازی فارسی این دیتاست شامل ۱۰۶۳ نمونه برای خلاصه‌سازی معنایی متون فارسی است. فرمت داده هر نمونه شامل سه فیلد است: instruction: دستور خلاصه‌سازی (بیش از ۴۰ دستور مختلف) input: متن ورودی (بین ۱۰۰ تا ۴۰۰ کلمه) output: خلاصه تولید شده (بین ۳۰ تا ۱۰۰ کلمه) تعداد نمونه‌ها ۱۰۶۳ نمونه با کیفیت بالا حوزه‌های پوشش داده شده اقتصادی (بانک، بورس، بیمه، تولید) سیاسی (داخلی، خارجی، مذاکرات) اجتماعی (آسیب‌شناسی، رفاه، اشتغال) علمی و… See the full description on the dataset page: https://huggingface.co/datasets/msdos34/persian-summary-dataset.texttext-generation1K<n<10K1 likes117 downloads15d agoHugging Face02yuyijiong /Book_Summary_Chinese 中文图书总结数据集 每个样本包含: 图书的一个章节、此章节的总结、图书名字,可以训练模型总结长文本的能力。数据主要来自较为著名的中文版小说。 texttext-generationn<1K27 likes86 downloads3y agoHugging Face03k007-a /story-summary Short Story Summarization Dataset Short stories from agentlans/euclaise-WritingPromptsX Summarized using Qwen/Qwen3.5-9B with the following prompt: Summarize the short story below in a single well-written paragraph that captures the main plot, key characters, central conflict, and resolution while preserving the original meaning and tone. Focus only on the most important details, avoid unnecessary specifics or minor subplots, and do not add interpretations or information that… See the full description on the dataset page: https://huggingface.co/datasets/k007-a/story-summary.texttext-generation10K<n<100K0 likes58 downloads27d agoHugging Face04agentlans /story-summary Short Story Summarization Dataset Short stories from agentlans/euclaise-WritingPromptsX Summarized using Qwen/Qwen3.5-9B with the following prompt: Summarize the short story below in a single well-written paragraph that captures the main plot, key characters, central conflict, and resolution while preserving the original meaning and tone. Focus only on the most important details, avoid unnecessary specifics or minor subplots, and do not add interpretations or information that… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/story-summary.texttext-generation10K<n<100K0 likes45 downloads3mo agoHugging Face05agentlans /summary-keywordstexttext-generation10K<n<100K0 likes40 downloads11mo agoHugging Face06agentlans /high-quality-summary Data from agentlans/high-quality-text sample_k10000 configuration Summaries generated using google/gemma-3-12b-it Summaries rewritten using agentlans/granite-3.3-2b-refiner Rewritten summaries checked against the original text using ibm-granite/granite-3.3-8b-instruct texttext-generation10K<n<100K0 likes37 downloads1y agoHugging Face07kilicai /turkish-doc-summary-review-analysis-30k-jsonl ⚠️ Superseded by v2 Bu v1 dataset'te exact duplicate yoktu; ancak belge ve cevap şablonları fazla tekrar ediyordu. Güncel v2 sürümünü kullanın: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl-v2 Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code:… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl.texttext-generation10K<n<100K0 likes23 downloads4mo agoHugging Face08kilicai /turkish-doc-summary-review-analysis-30k-jsonl-v2 Turkish Document Summary Review Analysis 30K JSONL v2 Tek dosya: train.jsonl. Bu v2 sürümü, v1'de görülen tekrar sorununu çözmek için yeniden üretildi: Daha fazla belge türü: proje önerisi, tutanak, denetim notu, şikâyet dosyası, politika taslağı, saha raporu, bütçe değerlendirmesi, risk kayıt formu, karar destek belgesi, olay inceleme raporu vb. Daha fazla alt görev: 30 farklı task_type. Exact duplicate + semantic template duplicate kontrolü. Cevap şablonları belgeye özel risk… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl-v2.texttext-generation10K<n<100K0 likes21 downloads4mo agoHugging Face09jungsanghyun /subtitle-summary-datasetgated Subtitle Summary & Keyword Dataset (YouTube) 국내 방송사 유튜브 자막 기반 스트리밍(누적) 요약 + 유튜브 검색 키워드 데이터셋. 방송을 5분 단위로 진행하며, 매 5분마다 (직전 누적 요약 + 새 5분 자막)으로 방송 대표 한 문장 요약을 갱신하고 검색 키워드를 생성합니다. 교사 모델 Qwen3-235B-A22B-Instruct-2507-FP8 (vLLM, 자막-only). HBKenerzai/LGUplus_summary_keyword와 동일 스키마. 규모 split 회차 레코드(5분) train 2,641 10,159 validation 294 971 합계 2,935 11,130 약 563시간 · 110개 채널 · 요약 중앙값 58자 · 도메인(유튜브 카테고리): Entertainment 2982 · Travel & Events 2081 · News &… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-dataset.tabularsummarization10K<n<100K0 likes21 downloads2mo agoHugging Face10Gladiaio /Instruct-Summary Dataset Card for "Instruct-Summary" This dataset is a combination of kmfoda/booksum, samsum, mosaicml/dolly_hhrlhf and yahma/alpaca-cleaned. textsummarization10K<n<100K3 likes18 downloads3y agoHugging Face11lianghsun /tw-law-context-summarygated Dataset Card for tw-law-context-summary 本資料集為中華民國(臺灣)法規條文之 LLM 摘要集,每筆樣本提供「法規名稱」對應的條列式摘要文字,可作為法規 RAG 系統的索引/簡介,或法律 chatbot 的初步說明資料。 Dataset Details Dataset Description 資料以法規為單位,由 LLM 對每部法規生成「立法依據、規範重點、施行日期、適用範圍、廢止狀態」等結構化摘要。每筆樣本欄位: text:摘要內容(條列式)。 name:法規名稱。 abandon_note:廢止/修訂註記,若為現行法規則為空字串。 token_count / word_count:保留為字串欄位(部分樣本為空)。 主要使用者為法規檢索、法律入門教育場景。 Curated by: Huang Liang Hsun Language(s) (NLP): Traditional Chinese License: cc-by-nc-sa-4.0… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-law-context-summary.texttext-generation10K<n<100K0 likes14 downloads5mo agoHugging Face12lianghsun /tw-processed-related-law-article-summarygated Dataset Card for tw-processed-related-law-article-summary 本資料集為中華民國(臺灣)法規「條文之間引用關係」的 LLM 摘要:每筆樣本針對某條條文與其引用之相關條文,提供結構化的關係分析(直接引用、技術延伸、規範一致性、實務操作、廢止情況等)。 Dataset Details Dataset Description tw-processed-related-law-article 已將每條條文與其引用之條文串接在一起;本資料集則進一步以 LLM 解析「該條與所引用條之間的關係」並輸出條列式摘要,內容通常包含: 直接引用與依據關係 技術規範的延伸 法規一致性 實務操作影響 廢止情況 可作為法規 RAG 系統的「條文關係層」索引,協助下游模型理解條文之間的依存。 Curated by: Huang Liang Hsun Language(s) (NLP): Traditional Chinese License: cc-by-nc-sa-4.0… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-related-law-article-summary.texttext-generation10K<n<100K0 likes13 downloads5mo agoHugging Face13silk-road /Haruhi-Dialogue-Speaker-Extract-And-Summary之前的 silk-road/Haruhi-Dialogue-Speaker-Extract 要求模型输出json格式,并且采取了CoT策略,感觉有一些难了 这一次把总结和抽取拆分成了两个任务 并且抽取的格式改为了csv格式。 texttext-generationn<1K0 likes12 downloads3y agoHugging Face14jungsanghyun /subtitle-summary-lgugated Broadcast Subtitle Summary & Keyword — (LGUplus source, our pipeline) HBKenerzai/LGUplus_summary_keyword의 방송 자막을 우리 형식으로 재구성한 뒤, **우리 파이프라인 (Qwen3-235B-A22B-Instruct-2507-FP8, vLLM, 자막-only)**으로 5분 단위 누적 요약 + 검색 키워드를 다시 생성한 데이터셋입니다. 원천 자막·메타 저작권은 원 출처(AI Hub / 방송사)에 있습니다. 규모 split 회차 레코드(5분) train 16,609 118,608 validation 1,846 13,177 합계 18,455 131,785 약 9,025시간 분량 · 요약 길이 중앙값 63자 · 도메인: 생활정보 9528 · 정보/토크 9458 · 퀴즈/게임 8123 · 푸드/요리 7221 · 토론/대담… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-lgu.tabularsummarization100K<n<1M0 likes12 downloads2mo agoHugging Face15jungsanghyun /subtitle-summary-sftgated Subtitle Summary & Keyword — SFT (YouTube) jungsanghyun/subtitle-summary-dataset를 SFT 학습용 messages 형식으로 가공. 형식 (system 없음, 단일 턴) 입력(user): [이전 요약] {last_summary}\n[자막] {5min_script} 출력(assistant): [요약] {한 문장}\n[키워드] {검색어} — 라벨 2줄(소형 모델 친화). split examples train 10,159 validation 971 chat_template.jinja({% generation %})로 assistant만 loss. 파싱 \[요약\]\s*(.+) / \[키워드\]\s*(.+). 라이선스 cc-by-nc-4.0, 연구·비상업. textsummarization10K<n<100K0 likes7 downloads2mo agoHugging Face16jungsanghyun /subtitle-summary-testsetgated Subtitle Summary & Keyword — Test set (YouTube) 방송 유튜브 자막 기반 누적 요약 + 검색 키워드 태스크의 평가용 테스트셋입니다. jungsanghyun/subtitle-summary-dataset의 train/validation과 겹치지 않는 별도 방송으로, 동일 파이프라인(Qwen3-235B, vLLM, 자막-only)으로 생성했습니다. 규모 split 회차 레코드(5분) test 114 519 약 25시간 · 요약 중앙값 59자 · 도메인: Entertainment 236 · News & Politics 103 · Pets & Animals 87 · Travel & Events 34 · Education 31 · Music 28 스키마 program_name, last_summary, 5min_script(입력) →… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-testset.tabularsummarizationn<1K0 likes7 downloads2mo agoHugging Face17ENERZAiKR /LGUplus_summary_keywordgated LGUplus Broadcast Streaming Summary & Keyword Dataset 한국어 방송 자막 기반 스트리밍(누적) 요약 + 키워드 추출 데이터셋입니다. 방송을 5분 단위로 진행하며, 매 5분마다 (직전까지의 요약 + 새 5분 자막)을 받아 방송 전체를 대표하는 한 문장 요약으로 갱신하고, 해당 5분의 핵심 키워드 1개를 뽑는 태스크를 위해 구축되었습니다. 데이터 구조 각 레코드 = 한 방송의 한 5분 구간(step). 필드 설명 program_name 프로그램명 last_summary 직전 시점까지의 누적 요약 (첫 청크는 "") — 입력 5min_script 이번 5분 구간 자막 ([역할] 발화 형식) — 입력 current_summary 갱신된 방송 전체 대표 요약 (한 문장) — 출력 current_5min_keyword 이번 5분 구간 핵심 키워드 (1개) — 출력… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/LGUplus_summary_keyword.tabularsummarization100K<n<1M0 likes6 downloads3mo agoHugging Face18jungsanghyun /subtitle-summary-lgu-sftgated Broadcast Subtitle Summary & Keyword — SFT (LGUplus source) jungsanghyun/subtitle-summary-lgu를 SFT 학습용 messages 형식으로 가공한 버전입니다. 형식 (system 없음, 단일 턴) 입력(user): [이전 요약] {last_summary}\n[자막] {5min_script} 출력(assistant): [요약] {한 문장}\n[키워드] {검색어} — 라벨 2줄(소형 모델 친화·파싱 용이). split examples train 118,608 validation 13,177 generation만 학습 chat_template.jinja({% generation %} 블록)로 assistant 응답만 loss. TRL SFTConfig(assistant_only_loss=True… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-lgu-sft.textsummarization100K<n<1M0 likes5 downloads2mo agoHugging Face19jungsanghyun /subtitle-summary-testset-sftgated Subtitle Summary & Keyword — Test set SFT (YouTube) jungsanghyun/subtitle-summary-testset를 messages 형식으로 가공한 평가용 버전. 형식 (system 없음, 단일 턴) 입력(user): [이전 요약] {last_summary}\n[자막] {5min_script} 출력(assistant): [요약] {한 문장}\n[키워드] {검색어} split examples test 519 파싱 \[요약\]\s*(.+) / \[키워드\]\s*(.+). chat_template.jinja 동봉. 라이선스 cc-by-nc-4.0, 연구·비상업. textsummarizationn<1K0 likes5 downloads2mo agoHugging Face20agentlans /high-quality-summary-v2 High Quality Long Text Summarization Dataset Input texts from agentlans/high-quality-text-long sample_k10000 config Summaries generated by google/gemma-3-12b-it Summaries rewritten by agentlans/granite-3.3-2b-reviser texttext-generation10K<n<100K2 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.