CoolFace
Datasetpublic

giuliadc/cnndm_5k

To create this dataset, the test split of CNN DAILYMAIL was filtered by using the code by Aumiller et al. (1) available at https://github.com/dennlinger/summaries/tree/main with following settings: min_length_summary = 18; min_length_reference = 250; length_metric = "whitespace" min_compression_ratio = 2.5 Furthermore: line breaks: every \n in the reference summaries (column "reference-summary") was replaced by a space. The articles (column "text") did not contain any line breaks non-breaking… See the full description on the dataset page: https://huggingface.co/datasets/giuliadc/cnndm_5k.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes6downloads
Dataset Card

To create this dataset, the test split of CNN DAILYMAIL was filtered by using the code by Aumiller et al. (1) available at https://github.com/dennlinger/summaries/tree/main with following settings: minlengthsummary = 18; minlengthreference = 250; lengthmetric = "whitespace" mincompression_ratio = 2.5

Furthermore:

  • line breaks: every \n in the reference summaries (column "reference-summary") was replaced by a space. The articles (column "text") did not contain any line breaks
  • non-breaking spaces: every \xa0 (both in articles and in summaries) was replaced by a space
  • bi-gramoverlapfraction between summary and original text <= 0.63, meaning that all summaries in the dataset are on the abstractive side

Then, 5k random rows were selected and kept in the dataset. All other rows were removed.