CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abisee /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.textsummarization100K<n<1M351 likes74k downloads3y agoHugging Face02gigant /cnn_dailymail_oreo_jinacolbertv2_10ktext10K<n<100K0 likes337 downloads2y agoHugging Face03imvladikon /he_cnn_dailymail Dataset Card for "he_cnn_dailymail" More Information needed textsummarization100K<n<1M0 likes333 downloads3y agoHugging Face04Subhav-K /cnn-dailymail-nemotron-embeddingstext100K<n<1M1 likes297 downloads2mo agoHugging Face05Subhav-K /cnn-dailymail-chunked-512-embeddingstext100K<n<1M2 likes213 downloads3mo agoHugging Face06extraordinarylab /cnn-dailymailtext100K<n<1M0 likes191 downloads11mo agoHugging Face07Gabriel /cnn_daily_swe Dataset Card for Swedish CNN Dailymail Dataset The Swedish CNN/DailyMail dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks. Dataset Summary Read about the full details at original English version: https://huggingface.co/datasets/cnn_dailymail Data Fields id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from article: a string containing the body of the news… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel/cnn_daily_swe.textsummarization100K<n<1M0 likes159 downloads4y agoHugging Face08Subhav-K /cnn-dailymail-bge-base-embeddingstext10K<n<100K0 likes147 downloads2mo agoHugging Face09cestwc /cnn_dailymail-snippetstext1M<n<10M0 likes142 downloads5y agoHugging Face10AmrahMaryam /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/AmrahMaryam/cnn_dailymail.textsummarization100K<n<1M0 likes140 downloads8mo agoHugging Face11marco-willi /cnnspot-smallimage100K<n<1M0 likes133 downloads8mo agoHugging Face12farzad0514 /summarized-hyperpartisan-news-by-facebook-bart-large-cnn-v1text100K<n<1M0 likes123 downloads1y agoHugging Face13argilla /cnn-dailymail-summaries Dataset Card for cnn-dailymail-summaries This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: cnn_daily_summaries.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/argilla/cnn-dailymail-summaries/raw/main/cnn_daily_summaries.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/argilla/cnn-dailymail-summaries.textsummarization100K<n<1M7 likes103 downloads2y agoHugging Face14cestwc /cnn_dailymail-coreftext100K<n<1M0 likes102 downloads4y agoHugging Face15whu9 /cnn_dailymail_ngrams_1_to_5 Dataset Card for "cnn_dailymail_ngrams_1_to_5" More Information needed text100M<n<1B0 likes100 downloads4y agoHugging Face16closji /seq2seq-cnndm Dataset Card for "seq2seq-cnndm" More Information needed text100K<n<1M0 likes85 downloads3y agoHugging Face17emirhanboge /cnn_dailymail_llama1btext100K<n<1M0 likes83 downloads2y agoHugging Face18ereverter /cnn_dailymail_extractive Data Card for Extractive CNN/DailyMail Dataset Overview This is an extractive version of the CNN/Dailymail dataset. The structure of this dataset is identical to the original except for a minor modification in the data representation and the introduction of labels to denote the extractive summary. The labels are generated following a greedy algorithm, as proposed by Liu (2019). The curation process can be found in the bertsum-hf repository. I am uploading it in case… See the full description on the dataset page: https://huggingface.co/datasets/ereverter/cnn_dailymail_extractive.textsummarization100K<n<1M6 likes79 downloads3y agoHugging Face19yakul259 /Stage_1_CNN_Saliencetext100K<n<1M0 likes77 downloads8mo agoHugging Face20nschantz21 /cnn_dailymail-parsedtext100K<n<1M0 likes76 downloads4y agoHugging Face21emirhanboge /squad_v2_codex_glue_cnn_dailymail_llama1b_modifiedtext100K<n<1M0 likes76 downloads2y agoHugging Face22celsowm /cnn_news_ptbr Dataset Card for "cnn_news_ptbr" More Information needed texttext-classification1K<n<10K3 likes74 downloads2y agoHugging Face23youssefkhalil320 /cnn_dailymail_t5_summariestext10K<n<100K0 likes70 downloads2y agoHugging Face24GlazJ /cn-news-impact-scores Chinese News Impact Scores 2024-2025 This dataset pairs complete Chinese financial-news collections for 2024 and 2025 with event-level market, industry/board, and stock impact scores. Data is stored in monthly Parquet shards. Dataset Structure raw_news: every collected news occurrence, including full text, a unique occurrence_id, and a stable news_id. impact_scores: one row per event-target pair with routing metadata and 16 impact dimensions. event_clusters:… See the full description on the dataset page: https://huggingface.co/datasets/GlazJ/cn-news-impact-scores.tabulartext-classification1M<n<10M0 likes70 downloads2mo agoHugging Face25cestwc /cnn_dailymail-metaeval100textn<1K0 likes62 downloads5y agoHugging Face26pszemraj /cnn_dailymail-cleaned cnn_dailymail: cleaned Original cnn_dailymail config 3.0.0 with the following changes: renamed columns to text and summary cleaning applied to summary column to correct punctuation, etc. textsummarization100K<n<1M0 likes62 downloads9mo agoHugging Face27cestwc /cnn_dailymail-test50textn<1K0 likes59 downloads5y agoHugging Face28oddadmix /aya_collection-translated_cnn_dailymail Dataset Card for "aya_collection-translated_cnn_dailymail" More Information needed text1M<n<10M0 likes58 downloads5mo agoHugging Face29Mithilss /cnn_dollybricks_platypus_bbq_2_0 Dataset Card for "cnn_dollybricks_platypus_bbq_2_0" More Information needed text10K<n<100K0 likes55 downloads3y agoHugging Face30Newvel /cnn_dailymail_sanitized_tokenizedtext100K<n<1M0 likes54 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.