datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cn_ner来源 https://github.com/liucongg/NLPDataSet
从网上收集数据,将CMeEE数据集、IMCS21_task1数据集、CCKS2017_task2数据集、CCKS2018_task1数据集、CCKS2019_task1数据集、CLUENER2020数据集、MSRA数据集、NLPCC2018_task4数据集、CCFBDCI数据集、MMC数据集、WanChuang数据集、PeopleDairy1998数据集、PeopleDairy2004数据集、GAIIC2022_task2数据集、WeiBo数据集、ECommerce数据集、FinanceSina数据集、BoSon数据集、Resume数据集、Bank数据集、FNED数据集和DLNER数据集等22个数据集进行整理清洗,构建一个较完善的中文NER数据集。
数据集清洗时,仅进行了简单地规则清洗,并将格式进行了统一化,标签为“BIO”。
处理后数据集详细信息,见数据集描述。
数据集由NJUST-TB一起整理。
由于部分数据包含嵌套实体的情况,所以转换成BIO标签时,长实体会覆盖短实体。
数据… See the full description on the dataset page: https://huggingface.co/datasets/ttxy/cn_ner.cnn-based-drowsiness-detection-data
CNN-Based Drowsiness Detection - Dataset
Preprocessed, auto-labeled face-crop images used to train the model in
notgoodkeeper/cnn-based-drowsiness-detection.
Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection
Collection
Frames were captured from a webcam, then run through:
Haar Cascade face detection -> crop + pad + resize to 412x412
MediaPipe Selfie Segmentation -> background replaced with white
CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.CNN_News_Articles_2011-2022
CNN News Articles 2011-2022 Dataset
Introduction
This dataset contains CNN News Articles from 2011 to 2022 after basic cleaning. The dataset includes the following information:
Category
Full text
The data was downloaded from Kaggle at this URL: https://www.kaggle.com/datasets/hadasu92/cnn-articles-after-basic-cleaning. The dataset was split into two sets:
Train set with 32,218 examples
Test set with 5,686 examples
Usage
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/CNN_News_Articles_2011-2022.NO-CNN-DailyMail
Dataset Card
Dataset Summary
NO-CNN-DailyMail is a Norwegian news summarization dataset partially machine translated from English version of CNN Dailymail Dataset. The summaries were written by journalists at CNN and the DailyMail. The dataset can be used for Machine reading comprehension and abstractive summarization tasks.
Data Instances
For each instance, there is an article string and a positive_sample string representing news article and abstractive… See the full description on the dataset page: https://huggingface.co/datasets/NorGLM/NO-CNN-DailyMail.CNN-Daily-Mail-Sinhala
Dataset Summary
This dataset card aims to be creating a new dataset or Sinhala news summarization tasks. It has been generated using [https://huggingface.co/datasets/cnn_dailymail] and google translate.
Data Instances
For each instance, there is a string for the article, a string for the highlights, and a string for the id. See the CNN / Daily Mail dataset viewer to explore more examples.
{'id': '0054d6d30dbcad772e20b22771153a2a9cbeaf62',
'article': '(CNN) -- An American… See the full description on the dataset page: https://huggingface.co/datasets/Hamza-Ziyard/CNN-Daily-Mail-Sinhala.CNN_NEWS_DATASETKaggle_CNN_Text_Summarizationcnn-summarization
cnn-summarization
1k rows randomly extracted from the cnn_dailymail dataset with reasoning traces generated by Qwen3-14b
Kaggle_CNN_Text_SummarizationCNN-english-newsChat-PCR_CNNDaily_clearcnn-hindiChat-PCR_CNNDaily_NoRatioPCR_CNNDailyCNN_Newscnn-daily-mailcnn_enrich_with_top_keywords
Dataset Card for processed_dataset_top.csv
This dataset is an enhanced version of the CNN/DailyMail summarization dataset. Articles have been preprocessed and keywords are prepended at the top of each article to provide additional context for fine-tuning summarization models.
Dataset Details
Dataset Description
The dataset includes news articles with keywords prepended at the top, formatted with special tokens for compatibility with transformer-based models.… See the full description on the dataset page: https://huggingface.co/datasets/VexPoli/cnn_enrich_with_top_keywords.Chat-PCR_CNNDailybitcoin_oversample_cnn_extractedCnnDailymailcnn-v2
Test train/val/test for cnn.
cnn_enrich_data_with_down_keywords
Dataset Card for processed_dataset_bottom.csv
This dataset is an enhanced version of the CNN/DailyMail summarization dataset. Articles have been preprocessed and keywords are appended at the bottom of each article to provide additional context for fine-tuning summarization models.
Dataset Details
Dataset Description
The dataset includes news articles with keywords appended at the bottom, formatted with special tokens for compatibility with transformer-based models. Keywords were… See the full description on the dataset page: https://huggingface.co/datasets/VexPoli/cnn_enrich_data_with_down_keywords.cnn-v4
Official train/val for cnn with sramble amr.
cnn_dailymail_nor500_cnn_test_data_hindiCnn-Hindi-Datasetcnndailymail_pair_generatedChat-PCR_CNNDaily_Paraphrasemnist_cnncnn-v3
Official train/val for cnn.
