CoolFace
Datasetpublic

navjordj/VG_summarization

VG Summarization Dataset The source of this dataset is Norsk Aviskorpus (Norwegian newspaper corpus). This corpus includes articles from Norway’s largest newspaper from 1998 to 2019. In this dataset, we used the first paragraph (lead) of each article as its summary. This dataset only includes articles from the Norwegian newspaper "VG". The quality of the summary-article pairs has not been evaluated. License Please refer to the license of Norsk Aviskorpus… See the full description on the dataset page: https://huggingface.co/datasets/navjordj/VG_summarization.

sourceHugging Faceupdated 3y agoView on Hugging Face
6likes484downloads
Dataset Card

VG Summarization Dataset

The source of this dataset is Norsk Aviskorpus (Norwegian newspaper corpus). This corpus includes articles from Norway’s largest newspaper from 1998 to 2019. In this dataset, we used the first paragraph (lead) of each article as its summary. This dataset only includes articles from the Norwegian newspaper "VG".

The quality of the summary-article pairs has not been evaluated.

License

Please refer to the license of Norsk Aviskorpus

Citation

If you are using this dataset in your work, please cite our master thesis which this dataset was a part of

@mastersthesis{navjord2023beyond,
  title={Beyond extractive: advancing abstractive automatic text summarization in Norwegian with transformers},
  author={Navjord, J{\o}rgen Johnsen and Korsvik, Jon-Mikkel Ryen},
  year={2023},
  school={Norwegian University of Life Sciences, {\AA}s}
}