datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RyokoAI_CNNovel125K
Dataset Card for CNNovel125K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.CNNovel125K
Dataset Card for CNNovel125K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/qqceqqq/CNNovel125K.CNNovel125K
Dataset Card for CNNovel125K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/beiwoshuisheng/CNNovel125K.cnn-dailymail-sinhala-continuous-pretrain
CNN DailyMail Sinhala Continuous Pretraining Dataset
Dataset Description
This dataset is designed for continuous pretraining of Sinhala Small Language Models (SLMs) and Large Language Models (LLMs).
The dataset was created by processing the original Sinhala news articles from:
CNN Daily Mail Sinhala Dataset
The article_sinhala field from the original dataset was extracted, cleaned, and concatenated into larger continuous text blocks suitable for language model… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/cnn-dailymail-sinhala-continuous-pretrain.cnn-dailymail-llama4-maverick-summary
CNN/DailyMail Summary Dataset (Llama-4-Maverick-17B-128E-Instruct-FP8)
Dataset Description
This dataset contains high-quality summaries of CNN and DailyMail news articles generated using the Llama-4-Maverick-17B-128E-Instruct-FP8 model. Each summary provides a concise, accurate overview of the main story while preserving key facts and context.
Dataset Features
High-quality summaries: Generated using Llama-4-Maverick-17B-128E-Instruct-FP8 model
Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/cnn-dailymail-llama4-maverick-summary.cnn_dailymail_tinyThis dataset is a subset of https://huggingface.co/datasets/cnn_dailymail.
The training set is composed of 2,000 examples of the original training set and the test set is composed of 1,000 examples of the original validation set.
We use the version 1.0.0 of the CNN/DailyMail dataset.
ai-vs-human-meta-llama-Llama-3.1-8B-Instruct-CNN
AI vs Human dataset on the CNN DailyNews
Dataset Description
This dataset showcases pairs of truncated text and their respective completions, crafted either by humans or an AI language model.
Each article was randomly truncated between 25% and 50% of its length.
The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation.
Data Fields
'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/ilyasoulk/ai-vs-human-meta-llama-Llama-3.1-8B-Instruct-CNN.Alpaca-cnn-dailymail
Data Summary
Data set Alpaca-cnn-dailymail is a data set version format changed by ccdv/cnn_dailymail to meet Alpaca fine-tuning Llama2. Only versions 3.0.0 and 2.0.0 were used for merging and as a key data set for the summary extraction task.
Licensing Information
The Alpaca-cnn-dailymail dataset version 1.0.0 is released under the Apache-2.0 License.
Citation Information
@inproceedings{see-etal-2017-get,
title = "Get To The Point: Summarization with… See the full description on the dataset page: https://huggingface.co/datasets/ZhongshengWang/Alpaca-cnn-dailymail.CnnDailymail
