datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-dialogue-to-soap-summary
Dataset Card for Synthetic Medical Dialogues and SOAP Summaries
Dataset Description
Abstract
This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/omi-health/medical-dialogue-to-soap-summary.tiny-random-model-summaryexcerpt-summary-longctx
Excerpt Summary (Long-Context)
Book-excerpt summarization at seven context lengths (2K – 256K tokens), built for
long-context supervised fine-tuning and context-length stress-testing. Each example
asks a model to summarize a passage within a target word count; the reference summary
was generated by an LLM (see Provenance).
Usage
from datasets import load_dataset
# config name = context length: "2k", "8k", "16k", "32k", "64k", "128k", "256k"
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lefft/excerpt-summary-longctx.Chinese-Patent-Summary高质量中文专利摘要数据集。
Chinese-Patent-Summary高质量中文专利摘要数据集。
bill_summary_us
Dataset Card for "bill_summary_us"
Dataset Summary
Dataset for summarization of summarization of US Congressional bills (bill_summary_us).
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in format(congress number + bill type + bill number + bill version).
congress: number of the congress.
bill_type: type of… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_summary_us.persian-summary-dataset
دیتاست خلاصهسازی فارسی
این دیتاست شامل ۱۰۶۳ نمونه برای خلاصهسازی معنایی متون فارسی است.
فرمت داده
هر نمونه شامل سه فیلد است:
instruction: دستور خلاصهسازی (بیش از ۴۰ دستور مختلف)
input: متن ورودی (بین ۱۰۰ تا ۴۰۰ کلمه)
output: خلاصه تولید شده (بین ۳۰ تا ۱۰۰ کلمه)
تعداد نمونهها
۱۰۶۳ نمونه با کیفیت بالا
حوزههای پوشش داده شده
اقتصادی (بانک، بورس، بیمه، تولید)
سیاسی (داخلی، خارجی، مذاکرات)
اجتماعی (آسیبشناسی، رفاه، اشتغال)
علمی و… See the full description on the dataset page: https://huggingface.co/datasets/msdos34/persian-summary-dataset.Book_Summary_Chinese
中文图书总结数据集
每个样本包含:
图书的一个章节、此章节的总结、图书名字,可以训练模型总结长文本的能力。数据主要来自较为著名的中文版小说。
synthetic-persian-chatbot-summary-retrieval
Dataset Summary
Synthetic Persian Chatbot Summary Retrieval (SynPerChatbotSumSRetrieval) is a Persian (Farsi) dataset designed for the novel Summary Retrieval task, introduced as part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using the GPT-4o-mini Large Language Model and is derived from the Synthetic Persian Chatbot Dataset. The core task is to retrieve the correct, pre-generated summary that corresponds to a given user–chatbot… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-summary-retrieval.synthetic-persian-chatbot-rag-summary-retrieval
Dataset Summary
Synthetic Persian Chatbot RAG Summary Retrieval (SynPerChatbotRAGSumSRetrieval) is a Persian (Farsi) dataset for the Summary Retrieval task, specifically built for Retrieval-Augmented Generation (RAG) systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the Synthetic Persian Chatbot RAG Dataset. It evaluates the ability of models to match conversations—possibly… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-summary-retrieval.ESFT-summarynews-summary
News Summary
The summary is translated to hindi using IndicTrans2.We additionally remove duplicates from the original dataset
Usage:Cross-lingual summarization
stark-summary
Dataset Card for Stark-Summary
🏠 Homepage | 💻 Github | 📄 Arxiv | 📕 PDF
List of Provided Model Series
Ultron-Summarizer-Series: 🤖 Ultron-Summarizer-1B | 🤖 Ultron-Summarizer-3B | 🤖 Ultron-Summarizer-8B
Ultron 7B: 🤖 Ultron-7B
🚨 Disclaimer: All models and datasets are intended for research purposes only.
Dataset Summary
Stark is a publicly available, large-scale, long-term multi-modal conversation dataset that encompasses a diverse range of social… See the full description on the dataset page: https://huggingface.co/datasets/passing2961/stark-summary.vietnamese-financial-summary
Vietnamese Financial News Summarization with Number Preservation
pi05-libero-plus-perturbation-summary
pi0.5 LIBERO-plus Perturbation Summary
This dataset stores the summary report for six completed LIBERO-plus perturbation evaluations of TensorAuto/tPi0.5-libero.
Files:
pi05_libero_plus_six_perturb_hf_report.md: human-readable report with Hugging Face links, success rates, and short analysis.
pi05_libero_plus_six_perturb_hf_report.json: machine-readable summary.
The failure-grid videos and per-category metadata are stored in the linked per-perturbation datasets.
story-summary
Short Story Summarization Dataset
Short stories from agentlans/euclaise-WritingPromptsX
Summarized using Qwen/Qwen3.5-9B with the following prompt:
Summarize the short story below in a single well-written paragraph that captures the main plot, key characters, central conflict, and resolution while preserving the original meaning and tone. Focus only on the most important details, avoid unnecessary specifics or minor subplots, and do not add interpretations or information that… See the full description on the dataset page: https://huggingface.co/datasets/k007-a/story-summary.story-summary
Short Story Summarization Dataset
Short stories from agentlans/euclaise-WritingPromptsX
Summarized using Qwen/Qwen3.5-9B with the following prompt:
Summarize the short story below in a single well-written paragraph that captures the main plot, key characters, central conflict, and resolution while preserving the original meaning and tone. Focus only on the most important details, avoid unnecessary specifics or minor subplots, and do not add interpretations or information that… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/story-summary.cnki_summaryTurkish-Knowledge-extraction_topic-summarythai-wiki-summary-dataset
Thai Wiki Summary Dataset
Rows: 3,000 rows (Cleaned)
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"หน่วยพื้นฐานในการแบ่งเขตแดนในโปแลนด์คือ เทศบาล (กมินา) เมืองก็เป็นเทศบาลด้วยเช่นกัน ทว่ามีตราตั้งให้เป็นเมือง ทั้งเมืองและเทศบาลปกครองโดยนายกเทศมนตรี ทว่าในเทศบาล นายกเทศมนตรีเรียกว่าโวกต์ ( วอยต์ในภาษาโปแลนด์) ส่วนในเมืองเรียกว่าเบอร์มิสตร์ ในเมืองใหญ่ ๆ บางเมืองมีความรับผิดชอบและอำนาจพิเศษ… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-wiki-summary-dataset.Longcontext-aozora-summary長文からの要約データセットです。
長文は以下の青空文庫データセットを利用しました。
globis-university/aozorabunko-clean
License
CC BY 4.0
high-quality-summary
Data from agentlans/high-quality-text sample_k10000 configuration
Summaries generated using google/gemma-3-12b-it
Summaries rewritten using agentlans/granite-3.3-2b-refiner
Rewritten summaries checked against the original text using ibm-granite/granite-3.3-8b-instruct
summary-keywordsncbi_genes_summarynepali_news_text_summaryvisual-novel-summary
Dataset Card for Visual Novel Summary
This dataset contains excerpts from visual novels, derived from the winglian/visual-novels-json dataset. Each excerpt is annotated with respect to elements of fiction, including setting, characters, plot, conflict, themes, point of view, and tone.
Dataset Details
Dataset Description
This dataset was created by splitting rows from the winglian/visual-novels-json dataset into chunks of 32 consecutive lines. Each chunk was… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/visual-novel-summary.Complaint_summary_dataset
Complaint Summary Dataset 📝⚖️
Complaint Summary Dataset is a high-quality collection of synthetic legal complaints and their concise summaries, designed for training AI models in legal incident summarization, text-to-text generation, and formal language understanding.
💡 Use Case
This dataset is perfect for:
Training LLMs (like LLaMA, Mistral) to summarize long legal complaint texts into precise summaries.
Fine-tuning legal assistants, legal chatbots, or documentation… See the full description on the dataset page: https://huggingface.co/datasets/navaneeth005/Complaint_summary_dataset.annoy-rawcode-1k-first200-summary
Annoy RawCode 1k First 200 Lines Summary
This dataset contains a summary of the first 200 lines of the rawcode_1k.jsonl file from the Annoy-DataSync repository.
Source
GitHub Repository: https://github.com/piekeniuszwu/Annoy-DataSync
Raw JSONL file: https://raw.githubusercontent.com/piekeniuszwu/Annoy-DataSync/main/data/rawcode_1k.jsonl
SHA-256 of first 200 lines: 98b90e39af38766c017adcea5d3a373d447d511637121b2e4424c3880df94a18
Summary
The summary JSON… See the full description on the dataset page: https://huggingface.co/datasets/dongbobo/annoy-rawcode-1k-first200-summary.nepali_news_text_summary_sharegptgrammar-summary
Overveiw of Data
Data is a collection of synthetic and open source content from Openstax. It is a combination of these two datasets, curated and segmented by grammar relationships in the meta. I used Mistral 7B to generate the metadata contain the grammar relationships.
ambrosfitz/cosmopedia_summary (14,000/25,000 selected)
ambrosfitz/10k_history_summary (10,000)
For a total of around 24,000 rows.
This dataset was used to train a T5 model on grammar attention when summarizing… See the full description on the dataset page: https://huggingface.co/datasets/ambrosfitz/grammar-summary.
