datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cnn_dailymail
Dataset Card for CNN Dailymail Dataset
Dataset Summary
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
Supported Tasks and Leaderboards
'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.cnn_dailymailCNN/DailyMail non-anonymized summarization dataset.
There are two features:
- article: text of news article, used as the document to be summarized
- highlights: joined text of highlights with <s> and </s> around each
highlight, which is the target summaryCNNDetectioncn_nameCNNovel125K
Dataset Card for CNNovel125K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other… See the full description on the dataset page: https://huggingface.co/datasets/RyokoAI/CNNovel125K.RyokoAI_CNNovel125K
Dataset Card for CNNovel125K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.cnn-dailymail-chunked-512-embeddingsCNNovel125K
Dataset Card for CNNovel125K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/qqceqqq/CNNovel125K.he_cnn_dailymail
Dataset Card for "he_cnn_dailymail"
More Information needed
cnn-dailymail-nemotron-embeddingscnn_dailymail_oreo_jinacolbertv2_10kcn_ner来源 https://github.com/liucongg/NLPDataSet
从网上收集数据,将CMeEE数据集、IMCS21_task1数据集、CCKS2017_task2数据集、CCKS2018_task1数据集、CCKS2019_task1数据集、CLUENER2020数据集、MSRA数据集、NLPCC2018_task4数据集、CCFBDCI数据集、MMC数据集、WanChuang数据集、PeopleDairy1998数据集、PeopleDairy2004数据集、GAIIC2022_task2数据集、WeiBo数据集、ECommerce数据集、FinanceSina数据集、BoSon数据集、Resume数据集、Bank数据集、FNED数据集和DLNER数据集等22个数据集进行整理清洗,构建一个较完善的中文NER数据集。
数据集清洗时,仅进行了简单地规则清洗,并将格式进行了统一化,标签为“BIO”。
处理后数据集详细信息,见数据集描述。
数据集由NJUST-TB一起整理。
由于部分数据包含嵌套实体的情况,所以转换成BIO标签时,长实体会覆盖短实体。
数据… See the full description on the dataset page: https://huggingface.co/datasets/ttxy/cn_ner.SVHN_CNN_Specialist_Zoocnnspot-smallatlas_free_cnn_datasetcnn_daily_swe
Dataset Card for Swedish CNN Dailymail Dataset
The Swedish CNN/DailyMail dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks.
Dataset Summary
Read about the full details at original English version: https://huggingface.co/datasets/cnn_dailymail
Data Fields
id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from
article: a string containing the body of the news… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel/cnn_daily_swe.actas-cnn-datasetcnn-dailymail-bge-base-embeddingsCNNovel125K
Dataset Card for CNNovel125K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/beiwoshuisheng/CNNovel125K.cnn_dailymail_nl This dataset is the CNN/Dailymail dataset translated to Dutch.
This is the original dataset:
```
load_dataset("cnn_dailymail", '3.0.0')
```
And this is the HuggingFace translation pipeline:
```
pipeline(
task='translation_en_to_nl',
model='Helsinki-NLP/opus-mt-en-nl',
tokenizer='Helsinki-NLP/opus-mt-en-nl')
```cnn_dailymail-snippetsgtzan-multi-cnncnn_dailymail
Dataset Card for CNN Dailymail Dataset
Dataset Summary
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
Supported Tasks and Leaderboards
'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/AmrahMaryam/cnn_dailymail.summarized-hyperpartisan-news-by-facebook-bart-large-cnn-v1cnn-dailymailcnn_muffins
CNN Muffins
A compact dog-versus-muffin image-classification dataset built around the
well-known visual confusion between Chihuahua faces and blueberry muffins.
Dataset structure
Split
Dogs
Muffins
Total
Train
319
161
480
Validation
36
18
54
Hard-16 benchmark
8
8
16
The hard-16 benchmark is isolated from train and validation. The JSONL files
use repository-relative image paths:
The benchmark labels follow the original 4x4 checkerboard layout… See the full description on the dataset page: https://huggingface.co/datasets/VatsaDev/cnn_muffins.cnn_dailymail
CNN_Dailymail
This repository hosts a copy of the CNN_Dailymail dataset, a large-scale dataset designed for evaluating abstractive text summarization systems.
CNN_Dailymail consists of news articles paired with human-written summaries, commonly used for training and evaluating models on summarization tasks. It contains articles from CNN and Daily Mail, covering a wide range of topics.
Contents
cnn_dailymail.jsonl (or your actual filename): the standard set of news… See the full description on the dataset page: https://huggingface.co/datasets/S3IC/cnn_dailymail.cnn-dailymail-summaries
Dataset Card for cnn-dailymail-summaries
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
cnn_daily_summaries.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/argilla/cnn-dailymail-summaries/raw/main/cnn_daily_summaries.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/argilla/cnn-dailymail-summaries.cnn_dailymail_ngrams_1_to_5
Dataset Card for "cnn_dailymail_ngrams_1_to_5"
More Information needed
cnn_dailymail-coref
