CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EdinburghNLP /xsum Dataset Card for "xsum" Dataset Summary Extreme Summarization (XSum) Dataset. There are three features: document: Input news article. summary: One sentence summary of the article. id: BBC ID of the article. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 257.30 MB Size of the generated dataset:… See the full description on the dataset page: https://huggingface.co/datasets/EdinburghNLP/xsum.textsummarization100K<n<1M152 likes16k downloads8mo agoHugging Face02GEM /xsumThis is the XSUM subset of the GEM benchmark.summarization3 likes1.5k downloads4y agoHugging Face03embedded-language-flows /xsum_validation_t50 likes923 downloads4mo agoHugging Face04VarunGumma /IGB_XSumtextsummarization10K<n<100K0 likes527 downloads1y agoHugging Face05embedded-language-flows /xsum_train_t5tabular100K<n<1M0 likes386 downloads4mo agoHugging Face06google-research-datasets /xsum_factualityNeural abstractive summarization models are highly prone to hallucinate content that is unfaithful to the input document. The popular metric such as ROUGE fails to show the severity of the problem. The dataset consists of faithfulness and factuality annotations of abstractive summaries for the XSum dataset. We have crowdsourced 3 judgements for each of 500 x 5 document-system pairs. This will be a valuable resource to the abstractive summarization community.summarization1K<n<10K6 likes247 downloads3y agoHugging Face07yairfeldman /xsumtext100K<n<1M0 likes153 downloads1y agoHugging Face08knkarthick /xsum Dataset Card for SAMSum Corpus Dataset Description Links Homepage: https://arxiv.org/abs/1808.08745 Repository: https://arxiv.org/abs/1808.08745 Paper: https://arxiv.org/abs/1808.08745 Point of Contact: https://huggingface.co/knkarthick Dataset Summary This repository contains data and code for our EMNLP 2018 paper "Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization".… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/xsum.textsummarization100K<n<1M6 likes150 downloads4y agoHugging Face09Gabriel /xsum_swe Dataset Card for Swedish Xsum Dataset The Swedish xsum dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks. Dataset Summary Read about the full details at original English version: https://huggingface.co/datasets/xsum Data Fields id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from document: a string containing the body of the news article summary: a string… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel/xsum_swe.tabularsummarization100K<n<1M0 likes144 downloads4y agoHugging Face10tkj000 /xsum_deepseek-moe-16b-chat_token_patternstext10K<n<100K0 likes142 downloads1y agoHugging Face11stacked-summaries /stacked-xsum xsum-stacked The current version (corresponding to the stacked-booksum release): v0.3. See the Stacked Summaries org page for what this is and why it exists. The maximum input length is 16384 tokens, and the maximum output length is 1024 tokens (measured with the Long-T5 tokenizer). stats [2023-01-09 19:36:25] INFO:root:INPUTS - basic stats - train [2023-01-09 19:36:26] INFO:root:{'num_columns': 5, 'num_rows': 204045, 'num_unique_target': 203107, 'num_unique_text':… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/stacked-xsum.tabularsummarization100K<n<1M2 likes125 downloads4y agoHugging Face12zuzannad1 /preprocessed_xsum Dataset Card for "preprocessed_xsum" More Information needed timeseries100K<n<1M0 likes120 downloads3y agoHugging Face13Eymen3455 /xsum_tr0 likes118 downloads6y agoHugging Face14mia-llm /xsum-MIA-Benchmarktext10K<n<100K0 likes98 downloads2y agoHugging Face15stacked-summaries /stacked-xsum-1024 stacked-xsum-1024 a "stacked" version of xsum Original Dataset: copy of the base dataset Stacked Rows: The original dataset is processed by stacking rows based on certain criteria: Maximum Input Length: The maximum length for input sequences is 1024 tokens in the longt5 model tokenizer. Maximum Output Length: The maximum length for output sequences is also 1024 tokens in the longt5 model tokenizer. Special Token: The dataset utilizes the [NEXT_CONCEPT] token to indicate a new… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/stacked-xsum-1024.tabularsummarization100K<n<1M1 likes93 downloads3y agoHugging Face16ExpertFlowPredictor /xsum_Qwen3-30B-A3B_moe_patternstext1K<n<10K0 likes91 downloads11mo agoHugging Face17bobox /xSum-processedtabular100K<n<1M0 likes80 downloads2y agoHugging Face18talgatzh /xsum-kk3Extreme Summarization (XSum) Dataset. There are three features: - document: Input news article. - summary: One sentence summary of the article. - id: BBC ID of the article.summarization100K<n<1M0 likes76 downloads2mo agoHugging Face19ExpertFlowPredictor /xsum_deepseek-moe-16b-chat_moe_patternstext1K<n<10K0 likes74 downloads11mo agoHugging Face20stacked-summaries /onlystacked-xsum-1024 stacked-summaries/onlystacked-xsum-1024 Same thing as stacked-summaries/stacked-xsum-1024 but filtered such that is_stacked=True. Please refer to the original dataset for info and to raise issues if needed. Basic info on train split: <class 'pandas.core.frame.DataFrame'> RangeIndex: 116994 entries, 0 to 116993 Data columns (total 6 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 document 116994 non-null string 1… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/onlystacked-xsum-1024.tabularsummarization100K<n<1M0 likes67 downloads3y agoHugging Face21tkj000 /xsum_Qwen3-30B-A3B_token_patternstext1K<n<10K0 likes67 downloads1y agoHugging Face22NbAiLab /norwegian-xsum5 likes66 downloads3y agoHugging Face23bertin-project /BOE-XSUM BOE-XSUM Balanced Dataset - Reviewed and Cleaned Description The BOE 2025 Dataset is a collection of BOE articles with extreme summaries of them. This dataset has been carefully balanced and cleaned to ensure its quality and usefulness in natural language processing (NLP) tasks, primarily for evaluating generative models. Read more in https://arxiv.org/abs/2509.24908 Dataset Content The dataset is composed of the following subsets (splits): train: Training… See the full description on the dataset page: https://huggingface.co/datasets/bertin-project/BOE-XSUM.textsummarization1K<n<10K0 likes65 downloads1y agoHugging Face24iohadrubin /mini_xsumtext10K<n<100K0 likes64 downloads4y agoHugging Face25ml6team /xsum_nl Dataset Card for XSum NL Dataset Summary This dataset is a machine translated dataset. It's the XSum dataset translated with this model from English to Dutch. See the Hugginface page of the original dataset for more information on the format of this dataset. Use with: from datasets import load_dataset load_dataset("csv", "ml6team/xsum_nl") Languages Dutch Dataset Structure Data Instances [More Information Needed] Data… See the full description on the dataset page: https://huggingface.co/datasets/ml6team/xsum_nl.3 likes61 downloads4y agoHugging Face26sentence-transformers /xsum Dataset Card for xsum This dataset is a collection of pairs of news articles and their summaries. See xsum for additional information. This dataset can be used directly with Sentence Transformers to train embedding models. Dataset Subsets pair subset Columns: "article", "summary" Column types: str, str Examples:{ 'article': 'Denny Solomona crossed for Castleford, but Wigan led at half-time against the run of play through Lewis Tierney\'s try and Matty… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/xsum.textfeature-extraction100K<n<1M0 likes51 downloads2y agoHugging Face27andreapdr /LID-XSUM LID-XSUM: Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors Datasets generated by aligning LLMs using Direct Preference Optimization to shift the machine-generated texts' (MGT) style toward human-written text (HWT). This dataset is intended to be used to augment the training set of documents to train more robust MGT detectors. Dataset Details The adversarial generations obtained in the paper "Stress-testing… See the full description on the dataset page: https://huggingface.co/datasets/andreapdr/LID-XSUM.text-classification3 likes51 downloads1y agoHugging Face28LM-Polygraph /xsum Dataset Card for xsum This is a preprocessed version of xsum dataset for benchmarks in LM-Polygraph. Dataset Details Dataset Description Curated by: https://huggingface.co/LM-Polygraph License: https://github.com/IINemo/lm-polygraph/blob/main/LICENSE.md Dataset Sources [optional] Repository: https://github.com/IINemo/lm-polygraph Uses Direct Use This dataset should be used for performing benchmarks on LM-polygraph.… See the full description on the dataset page: https://huggingface.co/datasets/LM-Polygraph/xsum.text100K<n<1M0 likes48 downloads1y agoHugging Face29Rexhaif /xsum_reducedtext100K<n<1M0 likes45 downloads4y agoHugging Face30mpalaval /scraped_xsum1textn<1K0 likes45 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.