datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cnn_dailymail
Dataset Card for CNN Dailymail Dataset
Dataset Summary
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
Supported Tasks and Leaderboards
'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.cnn_dailymail_oreo_jinacolbertv2_10khe_cnn_dailymail
Dataset Card for "he_cnn_dailymail"
More Information needed
cnn-dailymail-nemotron-embeddingscnn-dailymail-chunked-512-embeddingscnn-dailymailcnn_daily_swe
Dataset Card for Swedish CNN Dailymail Dataset
The Swedish CNN/DailyMail dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks.
Dataset Summary
Read about the full details at original English version: https://huggingface.co/datasets/cnn_dailymail
Data Fields
id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from
article: a string containing the body of the news… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel/cnn_daily_swe.cnn-dailymail-bge-base-embeddingscnn_dailymail-snippetscnn_dailymail
Dataset Card for CNN Dailymail Dataset
Dataset Summary
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
Supported Tasks and Leaderboards
'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/AmrahMaryam/cnn_dailymail.cnnspot-smallsummarized-hyperpartisan-news-by-facebook-bart-large-cnn-v1cnn-dailymail-summaries
Dataset Card for cnn-dailymail-summaries
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
cnn_daily_summaries.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/argilla/cnn-dailymail-summaries/raw/main/cnn_daily_summaries.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/argilla/cnn-dailymail-summaries.cnn_dailymail-corefcnn_dailymail_ngrams_1_to_5
Dataset Card for "cnn_dailymail_ngrams_1_to_5"
More Information needed
seq2seq-cnndm
Dataset Card for "seq2seq-cnndm"
More Information needed
cnn_dailymail_llama1bcnn_dailymail_extractive
Data Card for Extractive CNN/DailyMail Dataset
Overview
This is an extractive version of the CNN/Dailymail dataset. The structure of this dataset is identical to the original except for a minor modification in the data representation and the introduction of labels to denote the extractive summary.
The labels are generated following a greedy algorithm, as proposed by Liu (2019). The curation process can be found in the bertsum-hf repository. I am uploading it in case… See the full description on the dataset page: https://huggingface.co/datasets/ereverter/cnn_dailymail_extractive.Stage_1_CNN_Saliencecnn_dailymail-parsedsquad_v2_codex_glue_cnn_dailymail_llama1b_modifiedcnn_news_ptbr
Dataset Card for "cnn_news_ptbr"
More Information needed
cnn_dailymail_t5_summariescn-news-impact-scores
Chinese News Impact Scores 2024-2025
This dataset pairs complete Chinese financial-news collections for 2024 and 2025 with event-level market, industry/board, and stock impact scores. Data is stored in monthly Parquet shards.
Dataset Structure
raw_news: every collected news occurrence, including full text, a unique occurrence_id, and a stable news_id.
impact_scores: one row per event-target pair with routing metadata and 16 impact dimensions.
event_clusters:… See the full description on the dataset page: https://huggingface.co/datasets/GlazJ/cn-news-impact-scores.cnn_dailymail-metaeval100cnn_dailymail-cleaned
cnn_dailymail: cleaned
Original cnn_dailymail config 3.0.0 with the following changes:
renamed columns to text and summary
cleaning applied to summary column to correct punctuation, etc.
cnn_dailymail-test50aya_collection-translated_cnn_dailymail
Dataset Card for "aya_collection-translated_cnn_dailymail"
More Information needed
cnn_dollybricks_platypus_bbq_2_0
Dataset Card for "cnn_dollybricks_platypus_bbq_2_0"
More Information needed
cnn_dailymail_sanitized_tokenized
