CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abisee /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.textsummarization100K<n<1M351 likes75k downloads3y agoHugging Face02botp /RyokoAI_CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.texttext-classification1K<n<10K2 likes500 downloads3y agoHugging Face03qqceqqq /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/qqceqqq/CNNovel125K.texttext-classification10K<n<100K0 likes305 downloads6mo agoHugging Face04imvladikon /he_cnn_dailymail Dataset Card for "he_cnn_dailymail" More Information needed textsummarization100K<n<1M0 likes301 downloads3y agoHugging Face05Subhav-K /cnn-dailymail-chunked-512-embeddingstext100K<n<1M2 likes288 downloads3mo agoHugging Face06gigant /cnn_dailymail_oreo_jinacolbertv2_10ktext10K<n<100K0 likes272 downloads2y agoHugging Face07Subhav-K /cnn-dailymail-nemotron-embeddingstext100K<n<1M1 likes265 downloads2mo agoHugging Face08notgoodkeeper /cnn-based-drowsiness-detection-data CNN-Based Drowsiness Detection - Dataset Preprocessed, auto-labeled face-crop images used to train the model in notgoodkeeper/cnn-based-drowsiness-detection. Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection Collection Frames were captured from a webcam, then run through: Haar Cascade face detection -> crop + pad + resize to 412x412 MediaPipe Selfie Segmentation -> background replaced with white CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.imageimage-classification1K<n<10K1 likes191 downloads21d agoHugging Face09ttxy /cn_ner来源 https://github.com/liucongg/NLPDataSet 从网上收集数据,将CMeEE数据集、IMCS21_task1数据集、CCKS2017_task2数据集、CCKS2018_task1数据集、CCKS2019_task1数据集、CLUENER2020数据集、MSRA数据集、NLPCC2018_task4数据集、CCFBDCI数据集、MMC数据集、WanChuang数据集、PeopleDairy1998数据集、PeopleDairy2004数据集、GAIIC2022_task2数据集、WeiBo数据集、ECommerce数据集、FinanceSina数据集、BoSon数据集、Resume数据集、Bank数据集、FNED数据集和DLNER数据集等22个数据集进行整理清洗,构建一个较完善的中文NER数据集。 数据集清洗时,仅进行了简单地规则清洗,并将格式进行了统一化,标签为“BIO”。 处理后数据集详细信息,见数据集描述。 数据集由NJUST-TB一起整理。 由于部分数据包含嵌套实体的情况,所以转换成BIO标签时,长实体会覆盖短实体。 数据… See the full description on the dataset page: https://huggingface.co/datasets/ttxy/cn_ner.texttoken-classification100K<n<1M6 likes170 downloads3y agoHugging Face10extraordinarylab /cnn-dailymailtext100K<n<1M0 likes159 downloads11mo agoHugging Face11Gabriel /cnn_daily_swe Dataset Card for Swedish CNN Dailymail Dataset The Swedish CNN/DailyMail dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks. Dataset Summary Read about the full details at original English version: https://huggingface.co/datasets/cnn_dailymail Data Fields id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from article: a string containing the body of the news… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel/cnn_daily_swe.textsummarization100K<n<1M0 likes157 downloads4y agoHugging Face12Subhav-K /cnn-dailymail-bge-base-embeddingstext10K<n<100K0 likes150 downloads2mo agoHugging Face13beiwoshuisheng /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/beiwoshuisheng/CNNovel125K.texttext-classification10K<n<100K1 likes148 downloads8mo agoHugging Face14cestwc /cnn_dailymail-snippetstext1M<n<10M0 likes138 downloads5y agoHugging Face15marco-willi /cnnspot-smallimage100K<n<1M0 likes127 downloads8mo agoHugging Face16AmrahMaryam /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/AmrahMaryam/cnn_dailymail.textsummarization100K<n<1M0 likes124 downloads8mo agoHugging Face17farzad0514 /summarized-hyperpartisan-news-by-facebook-bart-large-cnn-v1text100K<n<1M0 likes123 downloads1y agoHugging Face18VatsaDev /cnn_muffins CNN Muffins A compact dog-versus-muffin image-classification dataset built around the well-known visual confusion between Chihuahua faces and blueberry muffins. Dataset structure Split Dogs Muffins Total Train 319 161 480 Validation 36 18 54 Hard-16 benchmark 8 8 16 The hard-16 benchmark is isolated from train and validation. The JSONL files use repository-relative image paths: The benchmark labels follow the original 4x4 checkerboard layout… See the full description on the dataset page: https://huggingface.co/datasets/VatsaDev/cnn_muffins.imageimage-classificationn<1K0 likes106 downloads21d agoHugging Face19S3IC /cnn_dailymail CNN_Dailymail This repository hosts a copy of the CNN_Dailymail dataset, a large-scale dataset designed for evaluating abstractive text summarization systems. CNN_Dailymail consists of news articles paired with human-written summaries, commonly used for training and evaluating models on summarization tasks. It contains articles from CNN and Daily Mail, covering a wide range of topics. Contents cnn_dailymail.jsonl (or your actual filename): the standard set of news… See the full description on the dataset page: https://huggingface.co/datasets/S3IC/cnn_dailymail.textsummarizationn<1K0 likes105 downloads9mo agoHugging Face20whu9 /cnn_dailymail_ngrams_1_to_5 Dataset Card for "cnn_dailymail_ngrams_1_to_5" More Information needed text100M<n<1B0 likes100 downloads4y agoHugging Face21argilla /cnn-dailymail-summaries Dataset Card for cnn-dailymail-summaries This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: cnn_daily_summaries.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/argilla/cnn-dailymail-summaries/raw/main/cnn_daily_summaries.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/argilla/cnn-dailymail-summaries.textsummarization100K<n<1M7 likes100 downloads2y agoHugging Face22cestwc /cnn_dailymail-coreftext100K<n<1M0 likes97 downloads4y agoHugging Face23mmichall /smclm-cnn-news@ARTICLE{11068992, author={Perełkiewicz, Michał and Dadas, Sławomir and Poświata, Rafał}, journal={IEEE Access}, title={SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation}, year={2025}, volume={}, number={}, pages={1-1}, keywords={Semantics;Transformers;Training;Encoding;Decoding;Syntactics;Measurement;Computational modeling;Autoencoders;Adaptation models;Autoregressive models;paraphrase generation;text embeddings;semantically… See the full description on the dataset page: https://huggingface.co/datasets/mmichall/smclm-cnn-news.text100K<n<1M0 likes85 downloads1y agoHugging Face24emirhanboge /cnn_dailymail_llama1btext100K<n<1M0 likes82 downloads2y agoHugging Face25closji /seq2seq-cnndm Dataset Card for "seq2seq-cnndm" More Information needed text100K<n<1M0 likes80 downloads3y agoHugging Face26GlazJ /cn-news-impact-scores Chinese News Impact Scores 2024-2025 This dataset pairs complete Chinese financial-news collections for 2024 and 2025 with event-level market, industry/board, and stock impact scores. Data is stored in monthly Parquet shards. Dataset Structure raw_news: every collected news occurrence, including full text, a unique occurrence_id, and a stable news_id. impact_scores: one row per event-target pair with routing metadata and 16 impact dimensions. event_clusters:… See the full description on the dataset page: https://huggingface.co/datasets/GlazJ/cn-news-impact-scores.tabulartext-classification1M<n<10M0 likes80 downloads2mo agoHugging Face27nschantz21 /cnn_dailymail-parsedtext100K<n<1M0 likes77 downloads4y agoHugging Face28ereverter /cnn_dailymail_extractive Data Card for Extractive CNN/DailyMail Dataset Overview This is an extractive version of the CNN/Dailymail dataset. The structure of this dataset is identical to the original except for a minor modification in the data representation and the introduction of labels to denote the extractive summary. The labels are generated following a greedy algorithm, as proposed by Liu (2019). The curation process can be found in the bertsum-hf repository. I am uploading it in case… See the full description on the dataset page: https://huggingface.co/datasets/ereverter/cnn_dailymail_extractive.textsummarization100K<n<1M6 likes77 downloads3y agoHugging Face29emirhanboge /squad_v2_codex_glue_cnn_dailymail_llama1b_modifiedtext100K<n<1M0 likes77 downloads2y agoHugging Face30yakul259 /Stage_1_CNN_Saliencetext100K<n<1M0 likes76 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.