CoolFace
Datasetpublic

CxsGHost/CNNSum

CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels Paper     GitHub [2025.5] - Accepted to Findings of ACL 2025 [2025.1] - Add inference script [2024.12] - CNNSum Dataset Release We are excited to announce the release of the CNNSum dataset! As outlined in Section 3.1 and Appendix E of our paper, we have conducted a final round of manual cleaning to address any possible omissions.… See the full description on the dataset page: https://huggingface.co/datasets/CxsGHost/CNNSum.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
5likes59downloads
Dataset Card

CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels

Paper     GitHub

[2025.5] - Accepted to Findings of ACL 2025
[2025.1] - Add inference script
[2024.12] - CNNSum Dataset Release

We are excited to announce the release of the CNNSum dataset!

As outlined in Section 3.1 and Appendix E of our paper, we have conducted a final round of manual cleaning to address any possible omissions. This process affects only a minimal number of samples, ensuring that the length statistics reported in our paper remain virtually unchanged.


Dataset Details

The dataset consists of two primary fields:

  • context: The novel excerpts
  • summary: The manually annotated summaries