CoolFace
Datasetpublic

Riksarkivet/mini_cleaned_diachronic_swe

Dataset Card for mini_cleaned_diachronic_swe The Swedish Diachronic Corpus is a project funded by Swe-Clarin and provides a corpus of texts covering the time period from Old Swedish. The dataset has been preprocessed and can be recreated from here: Src_code. Dataset Summary The dataset has been filtered with the metadata: Manueally transcribed or post-ocr correction No scrambled sentences Year of origin: 15-19th centuary Data Splits This will be… See the full description on the dataset page: https://huggingface.co/datasets/Riksarkivet/mini_cleaned_diachronic_swe.

sourceHugging Facemitupdated 4y agoView on Hugging Face
0likes102downloads
Dataset Card

Dataset Card for minicleaneddiachronic_swe

The Swedish Diachronic Corpus is a project funded by Swe-Clarin and provides a corpus of texts covering the time period from Old Swedish. The dataset has been preprocessed and can be recreated from here: Src_code.

Dataset Summary

The dataset has been filtered with the metadata:

  • —Manueally transcribed or post-ocr correction
  • —No scrambled sentences
  • —Year of origin: 15-19th centuary

Data Splits

This will be further extended!

Dataset SplitNumber of Instances in Split
Train352137
Test7187

Acknowledgements

We gratefully acknowledge SWE-clarin for the datasets.

Citation Information

Eva Pettersson and Lars Borin (2022) Swedish Diachronic Corpus In Darja Fišer & Andreas Witt (eds.), CLARIN. The Infrastructure for Language Resources. Berlin: deGruyter. https://degruyter.com/document/doi/10.1515/9783110767377-022/html