CoolFace
Datasetpublic

giuliadc/cnndm_5k

To create this dataset, the test split of CNN DAILYMAIL was filtered by using the code by Aumiller et al. (1) available at https://github.com/dennlinger/summaries/tree/main with following settings: min_length_summary = 18; min_length_reference = 250; length_metric = "whitespace" min_compression_ratio = 2.5 Furthermore: line breaks: every \n in the reference summaries (column "reference-summary") was replaced by a space. The articles (column "text") did not contain any line breaks non-breaking… See the full description on the dataset page: https://huggingface.co/datasets/giuliadc/cnndm_5k.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes7downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
giuliadc/cnndm_5k · CoolFace