Riksarkivet/mini_cleaned_diachronic_swe
Dataset Card for mini_cleaned_diachronic_swe The Swedish Diachronic Corpus is a project funded by Swe-Clarin and provides a corpus of texts covering the time period from Old Swedish. The dataset has been preprocessed and can be recreated from here: Src_code. Dataset Summary The dataset has been filtered with the metadata: Manueally transcribed or post-ocr correction No scrambled sentences Year of origin: 15-19th centuary Data Splits This will be… See the full description on the dataset page: https://huggingface.co/datasets/Riksarkivet/mini_cleaned_diachronic_swe.
Dataset Card for minicleaneddiachronic_swe
The Swedish Diachronic Corpus is a project funded by Swe-Clarin and provides a corpus of texts covering the time period from Old Swedish. The dataset has been preprocessed and can be recreated from here: Src_code.
Dataset Summary
The dataset has been filtered with the metadata:
- Manueally transcribed or post-ocr correction
- No scrambled sentences
- Year of origin: 15-19th centuary
Data Splits
This will be further extended!
Acknowledgements
We gratefully acknowledge SWE-clarin for the datasets.
Citation Information
Eva Pettersson and Lars Borin (2022) Swedish Diachronic Corpus In Darja Fišer & Andreas Witt (eds.), CLARIN. The Infrastructure for Language Resources. Berlin: deGruyter. https://degruyter.com/document/doi/10.1515/9783110767377-022/html
