Riksarkivet/mini_cleaned_diachronic_swe
Dataset Card for mini_cleaned_diachronic_swe The Swedish Diachronic Corpus is a project funded by Swe-Clarin and provides a corpus of texts covering the time period from Old Swedish. The dataset has been preprocessed and can be recreated from here: Src_code. Dataset Summary The dataset has been filtered with the metadata: Manueally transcribed or post-ocr correction No scrambled sentences Year of origin: 15-19th centuary Data Splits This will be… See the full description on the dataset page: https://huggingface.co/datasets/Riksarkivet/mini_cleaned_diachronic_swe.
0103
