datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xsum
Dataset Card for SAMSum Corpus
Dataset Description
Links
Homepage: https://arxiv.org/abs/1808.08745
Repository: https://arxiv.org/abs/1808.08745
Paper: https://arxiv.org/abs/1808.08745
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
This repository contains data and code for our EMNLP 2018 paper "Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization".… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/xsum.deepsynth-en-xsum
DeepSynth - XSum BBC News Summarization
Dataset Description
BBC news articles with single-sentence summaries. Focused on extreme summarization where the summary is
a single sentence capturing the essence of the article.
This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into images and processed through a frozen 380M parameter visual encoder… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-en-xsum.xsum-llama4-maverick-summary
XSum Summary Dataset (Llama-4-Maverick-17B-128E-Instruct-FP8)
Dataset Description
This dataset contains high-quality summaries of BBC news articles from the XSum (Extreme Summarization) dataset, generated using the Llama-4-Maverick-17B-128E-Instruct-FP8 model. Each summary provides a concise, accurate overview of the main story while preserving key facts and context.
Dataset Features
High-quality summaries: Generated using… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/xsum-llama4-maverick-summary.xsum_tinyThis dataset is a subset of https://huggingface.co/datasets/EdinburghNLP/xsum.
The training set is composed of 2,000 examples of the original training set and the test set is composed of 1,000 examples of the original validation set.
