datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
news-summary
Dataset Card for "news-summary"
Dataset Summary
Officially it was supposed to be used for classification but, can you use this data set to summarize news articles?
Languages
english
Citation Information
Acknowledgements
Ahmed H, Traore I, Saad S. “Detecting opinion spams and fake news using text classification”, Journal of Security and Privacy, Volume 1, Issue 1, Wiley, January/February 2018.
Ahmed H, Traore I, Saad S. (2017) “Detection of Online… See the full description on the dataset page: https://huggingface.co/datasets/argilla/news-summary.hf_paper_summary
Paper Espresso Dataset
This dataset repository contains structured metadata, summaries, and topical analysis for trending AI research papers, as presented in the paper Paper Espresso: From Paper Overload to Research Insight.
Paper Espresso is an open-source platform designed to automatically discover, summarize, and analyze trending research papers from arXiv. The system uses large language models (LLMs) to generate structured summaries, topical labels, and keywords. Over 35… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/hf_paper_summary.asta-summary-citation-counts
Dataset Summary
This dataset tracks which scientific papers are most often cited by Asta, an agentic research platform that uses retrieval-augmented generation (RAG) to answer scientific questions. Each record is a paper cited by Asta's Summarize Literature tool, ranked by the number of times the system cited that paper. Across more than 113,000 user queries, we track 4M citations to over 2M distinct papers. By making this data public, we aim to create a transparent, trackable… See the full description on the dataset page: https://huggingface.co/datasets/allenai/asta-summary-citation-counts.medical-dialogue-to-soap-summary
Dataset Card for Synthetic Medical Dialogues and SOAP Summaries
Dataset Description
Abstract
This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/omi-health/medical-dialogue-to-soap-summary.rtpurbo-block-summary-failed-experiment-artifacts
RTPurbo block-summary failed experiment artifacts
Immutable research artifacts from the entropy-calibrated 64-token block-summary investigation for RTPurbo/Qwen3.5-0.8B.
The tested static tangent and CAMS geometries failed the registered selector fidelity/traffic gate. This repository preserves the reusable feature/teacher caches, schedules, checkpoints, controls, scoreboards, and diagnostic evidence needed to reproduce or revisit that conclusion. It is an experiment archive… See the full description on the dataset page: https://huggingface.co/datasets/danym/rtpurbo-block-summary-failed-experiment-artifacts.pn_summaryA well-structured summarization dataset for the Persian language consists of 93,207 records. It is prepared for Abstractive/Extractive tasks (like cnn_dailymail for English). It can also be used in other scopes like Text Generation, Title Generation, and News Category Classification.
It is imperative to consider that the newlines were replaced with the `[n]` symbol. Please interpret them into normal newlines (for ex. `t.replace("[n]", "\n")`) and then use them for your purposes.clinvar_variant_summarysource data from https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz
arxiv_summarytutorials_summary
Tutorials Summary Text Dataset
This is the summary text dataset of sysmlv2's official tutorials pdf. With the text explanation and code examples in each page, organized in both Chinese and English natural language text. Useful for training LLM and teach it the basic knowledge and conceptions of sysmlv2.
182 records in total.
English Full Summary
page_1-41.md
page_42-81.md
page_82-121.md
page_122-161.md
page_162-183.md
中文完整版
page_1-41.md
page_42-81.md
page_82-121.md… See the full description on the dataset page: https://huggingface.co/datasets/sysmlv2research/tutorials_summary.terminal-wm-sft-summary-v2
Terminal World-Model SFT (summary, corrected v2)
Corrected supervised fine-tuning data derived from
open-thoughts/OpenThoughts-Agent-SFT-100K
Terminus traces. Schema version: v2-observation-action.
Causal format
Every world-model transition is serialized as:
system: target-specific world-model instruction
user: task context (turn 1) + real current observation_t + executed action_t
assistant: target derived from the real observation_t+1
Rows remain multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/terminal-wm-sft-summary-v2.bbc-news-summary
About Dataset
Context
Text summarization is a way to condense the large amount of information into a concise form by the process of selection of important information and discarding unimportant and redundant information. With the amount of textual information present in the world wide web the area of text summarization is becoming very important. The extractive summarization is the one where the exact sentences present in the document are used as summaries. The extractive… See the full description on the dataset page: https://huggingface.co/datasets/gopalkalpande/bbc-news-summary.TextCaps-Caption-Summary
Description
Multiple Captions of TextCaps dataset summarized into one using slauw87/bart_summarisation BART model.
tiny-random-model-summaryearnings-call-llama4-maverick-summary
Earnings Call Summary Dataset (Llama-4-Maverick-17B-128E-Instruct-FP8)
Dataset Description
This dataset contains comprehensive summaries of corporate earnings call transcripts generated using the Llama-4-Maverick-17B-128E-Instruct-FP8 model. Each summary provides structured insights into company performance, strategic initiatives, market conditions, and forward-looking guidance.
Dataset Features
High-quality summaries: Generated using… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/earnings-call-llama4-maverick-summary.MeetingBank-QA-Summary
Dataset Card for MeetingBank-QA-Summary
This dataset is introduced in LLMLingua-2 (Pan et al., 2024) and is designed to assess the performance of compressed meeting transcripts on downstream tasks such as question answering (QA) and summarization.
It includes 862 meeting transcripts from the test set of meeting transcripts introduced in MeetingBank (Hu et al, 2023) as the context, togeter with QA pairs and summaries that were generated by GPT-4 for each context transcripts.… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MeetingBank-QA-Summary.excerpt-summary-longctx
Excerpt Summary (Long-Context)
Book-excerpt summarization at seven context lengths (2K – 256K tokens), built for
long-context supervised fine-tuning and context-length stress-testing. Each example
asks a model to summarize a passage within a target word count; the reference summary
was generated by an LLM (see Provenance).
Usage
from datasets import load_dataset
# config name = context length: "2k", "8k", "16k", "32k", "64k", "128k", "256k"
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lefft/excerpt-summary-longctx.huggingthreat-secret-loyalties-summary
Secret Loyalties evaluation summary
Aggregate behavioral evidence from 1,564 generations across three gated Secret
Loyalties model organisms and the matched Qwen baseline. This public dataset
contains no prompts, raw generations, tokens, private contact information, or
gated weights.
What was observed
Model
Entity preference
Principal swap
Refusal probe
Self-report
Alamerton/sl-organism-a-7b
Companies inconclusive due to 38% order sensitivity; people… See the full description on the dataset page: https://huggingface.co/datasets/LihiShalmon/huggingthreat-secret-loyalties-summary.medical-dialogue-to-soap-summary-3kwiki_summary\
The dataset extracted from Persian Wikipedia into the form of articles and highlights and cleaned the dataset into pairs of articles and highlights and reduced the articles' length (only version 1.0.0) and highlights' length to a maximum of 512 and 128, respectively, suitable for parsBERT.korean-medical-dialogue-summary-datasetsummary微调google/mt5-base模型,做文章摘要
import torch
from transformers import T5ForConditionalGeneration, T5Tokenizer
model_path = "twwch/mt5-base-summary"
model = T5ForConditionalGeneration.from_pretrained(model_path)
tokenizer = T5Tokenizer.from_pretrained(model_path)
device = torch.device('cuda:0') if torch.cuda.is_available() else torch.device('cpu')
model.to(device)
model.eval()
text = """
什么是Nginx… See the full description on the dataset page: https://huggingface.co/datasets/twwch/summary.Chinese-Patent-Summary高质量中文专利摘要数据集。
leader-training-tokenized-fixed-summary-16k-fulltb2-eval-qwen3-8b-wm-summary-klanchor
Terminal-Bench 2.0 eval results - Qwen3-8B WM summary KL-anchor SFT
Compact Harbor eval results for violetxi/qwen3-8b-terminal-wm-summary-klanchor on
Terminal-Bench 2.0 using the terminus-2 agent harness.
Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward,
per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata.
Raw terminal recordings, panes, and completion logs are not included.… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb2-eval-qwen3-8b-wm-summary-klanchor.OCT-summary-Datasettb2-eval-qwen3-8b-wm-summary-mixed-clean-vista
Terminal-Bench 2.0 eval results - Qwen3-8B WM summary mixed clean Vista SFT
Compact Harbor eval results for violetxi/qwen3-8b-terminal-wm-summary-mixed-clean-vista on
Terminal-Bench 2.0 using the terminus-2 agent harness.
Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward,
per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata.
Raw terminal recordings, panes, and completion logs are… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb2-eval-qwen3-8b-wm-summary-mixed-clean-vista.arm-summaryannotations_creators:
other
language_creators:
other
languages:
hy-AM
licenses:
unknown
multilinguality:
monolingual
pretty_name: arm-sum
size_categories:
unknown
source_datasets:
original
task_categories:
conditional-text-generation
task_ids:
summarization
Chinese-Patent-Summary高质量中文专利摘要数据集。
QMugs_Summaryafricanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
