CoolFace
Datasetpublic

Gabriel/cnn_daily_swe

Dataset Card for Swedish CNN Dailymail Dataset The Swedish CNN/DailyMail dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks. Dataset Summary Read about the full details at original English version: https://huggingface.co/datasets/cnn_dailymail Data Fields id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from article: a string containing the body of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel/cnn_daily_swe.

sourceHugging Facemitupdated 4y agoView on Hugging Face
0likes156downloads
Dataset Card

Dataset Card for Swedish CNN Dailymail Dataset

The Swedish CNN/DailyMail dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks.

Dataset Summary

Read about the full details at original English version: https://huggingface.co/datasets/cnn_dailymail

Data Fields

  • id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from
  • article: a string containing the body of the news article
  • highlights: a string containing the highlight of the article as written by the article author

Data Splits

The Swedish CNN/DailyMail dataset follows the same splits as the original English version and has 3 splits: train, validation, and test.

Dataset SplitNumber of Instances in Split
Train287,113
Validation13,368
Test11,490