CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abisee /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.textsummarization100K<n<1M351 likes75k downloads3y agoHugging Face02ccdv /cnn_dailymailCNN/DailyMail non-anonymized summarization dataset. There are two features: - article: text of news article, used as the document to be summarized - highlights: joined text of highlights with <s> and </s> around each highlight, which is the target summarysummarization100K<n<1M35 likes30k downloads4y agoHugging Face03sywang /CNNDetection6 likes4.1k downloads2y agoHugging Face04raymondt /cn_name0 likes2k downloads9mo agoHugging Face05RyokoAI /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other… See the full description on the dataset page: https://huggingface.co/datasets/RyokoAI/CNNovel125K.text-classification100K<n<1M30 likes1.7k downloads3y agoHugging Face06botp /RyokoAI_CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.texttext-classification1K<n<10K2 likes501 downloads3y agoHugging Face07Subhav-K /cnn-dailymail-chunked-512-embeddingstext100K<n<1M2 likes311 downloads3mo agoHugging Face08qqceqqq /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/qqceqqq/CNNovel125K.texttext-classification10K<n<100K0 likes289 downloads6mo agoHugging Face09imvladikon /he_cnn_dailymail Dataset Card for "he_cnn_dailymail" More Information needed textsummarization100K<n<1M0 likes285 downloads3y agoHugging Face10Subhav-K /cnn-dailymail-nemotron-embeddingstext100K<n<1M1 likes274 downloads2mo agoHugging Face11gigant /cnn_dailymail_oreo_jinacolbertv2_10ktext10K<n<100K0 likes272 downloads2y agoHugging Face12ttxy /cn_ner来源 https://github.com/liucongg/NLPDataSet 从网上收集数据,将CMeEE数据集、IMCS21_task1数据集、CCKS2017_task2数据集、CCKS2018_task1数据集、CCKS2019_task1数据集、CLUENER2020数据集、MSRA数据集、NLPCC2018_task4数据集、CCFBDCI数据集、MMC数据集、WanChuang数据集、PeopleDairy1998数据集、PeopleDairy2004数据集、GAIIC2022_task2数据集、WeiBo数据集、ECommerce数据集、FinanceSina数据集、BoSon数据集、Resume数据集、Bank数据集、FNED数据集和DLNER数据集等22个数据集进行整理清洗,构建一个较完善的中文NER数据集。 数据集清洗时,仅进行了简单地规则清洗,并将格式进行了统一化,标签为“BIO”。 处理后数据集详细信息,见数据集描述。 数据集由NJUST-TB一起整理。 由于部分数据包含嵌套实体的情况,所以转换成BIO标签时,长实体会覆盖短实体。 数据… See the full description on the dataset page: https://huggingface.co/datasets/ttxy/cn_ner.texttoken-classification100K<n<1M6 likes195 downloads3y agoHugging Face13zerrougi /SVHN_CNN_Specialist_Zoo0 likes186 downloads4d agoHugging Face14marco-willi /cnnspot-smallimage100K<n<1M0 likes169 downloads8mo agoHugging Face15neurovlm /atlas_free_cnn_dataset0 likes165 downloads3mo agoHugging Face16Gabriel /cnn_daily_swe Dataset Card for Swedish CNN Dailymail Dataset The Swedish CNN/DailyMail dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks. Dataset Summary Read about the full details at original English version: https://huggingface.co/datasets/cnn_dailymail Data Fields id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from article: a string containing the body of the news… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel/cnn_daily_swe.textsummarization100K<n<1M0 likes156 downloads4y agoHugging Face17f3r21 /actas-cnn-datasetdocument0 likes152 downloads3mo agoHugging Face18Subhav-K /cnn-dailymail-bge-base-embeddingstext10K<n<100K0 likes150 downloads2mo agoHugging Face19beiwoshuisheng /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/beiwoshuisheng/CNNovel125K.texttext-classification10K<n<100K1 likes148 downloads8mo agoHugging Face20ml6team /cnn_dailymail_nl This dataset is the CNN/Dailymail dataset translated to Dutch. This is the original dataset: ``` load_dataset("cnn_dailymail", '3.0.0') ``` And this is the HuggingFace translation pipeline: ``` pipeline( task='translation_en_to_nl', model='Helsinki-NLP/opus-mt-en-nl', tokenizer='Helsinki-NLP/opus-mt-en-nl') ```100K<n<1M14 likes147 downloads4y agoHugging Face21cestwc /cnn_dailymail-snippetstext1M<n<10M0 likes142 downloads5y agoHugging Face22tristantanjh /gtzan-multi-cnnimage0 likes129 downloads5mo agoHugging Face23AmrahMaryam /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/AmrahMaryam/cnn_dailymail.textsummarization100K<n<1M0 likes124 downloads8mo agoHugging Face24farzad0514 /summarized-hyperpartisan-news-by-facebook-bart-large-cnn-v1text100K<n<1M0 likes123 downloads1y agoHugging Face25extraordinarylab /cnn-dailymailtext100K<n<1M0 likes113 downloads11mo agoHugging Face26VatsaDev /cnn_muffins CNN Muffins A compact dog-versus-muffin image-classification dataset built around the well-known visual confusion between Chihuahua faces and blueberry muffins. Dataset structure Split Dogs Muffins Total Train 319 161 480 Validation 36 18 54 Hard-16 benchmark 8 8 16 The hard-16 benchmark is isolated from train and validation. The JSONL files use repository-relative image paths: The benchmark labels follow the original 4x4 checkerboard layout… See the full description on the dataset page: https://huggingface.co/datasets/VatsaDev/cnn_muffins.imageimage-classificationn<1K0 likes106 downloads19d agoHugging Face27S3IC /cnn_dailymail CNN_Dailymail This repository hosts a copy of the CNN_Dailymail dataset, a large-scale dataset designed for evaluating abstractive text summarization systems. CNN_Dailymail consists of news articles paired with human-written summaries, commonly used for training and evaluating models on summarization tasks. It contains articles from CNN and Daily Mail, covering a wide range of topics. Contents cnn_dailymail.jsonl (or your actual filename): the standard set of news… See the full description on the dataset page: https://huggingface.co/datasets/S3IC/cnn_dailymail.textsummarizationn<1K0 likes102 downloads9mo agoHugging Face28argilla /cnn-dailymail-summaries Dataset Card for cnn-dailymail-summaries This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: cnn_daily_summaries.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/argilla/cnn-dailymail-summaries/raw/main/cnn_daily_summaries.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/argilla/cnn-dailymail-summaries.textsummarization100K<n<1M7 likes101 downloads2y agoHugging Face29whu9 /cnn_dailymail_ngrams_1_to_5 Dataset Card for "cnn_dailymail_ngrams_1_to_5" More Information needed text100M<n<1B0 likes99 downloads4y agoHugging Face30cestwc /cnn_dailymail-coreftext100K<n<1M0 likes97 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.