k8mpass-fraktur-series/fraktur-baltic-corpus
Fraktur Baltic Corpus Fraktur Baltic Corpus is a multilingual dataset based on historical German-language books printed in the Russian Empire during the 18th–19th centuries, primarily in Fraktur typeface. Each entry in the dataset contains: Raw OCR text from historical Fraktur sources Normalized German version Translations into seven languages: English, Russian, Estonian, Swedish, Finnish, Danish, and Modern German Volume 1: Hansen, Geschichte der Stadt Narva… See the full description on the dataset page: https://huggingface.co/datasets/k8mpass-fraktur-series/fraktur-baltic-corpus.
Fraktur Baltic Corpus
Fraktur Baltic Corpus is a multilingual dataset based on historical German-language books printed in the Russian Empire during the 18th–19th centuries, primarily in Fraktur typeface.
Each entry in the dataset contains:
- Raw OCR text from historical Fraktur sources
- Normalized German version
- Translations into seven languages: English, Russian, Estonian, Swedish, Finnish, Danish, and Modern German
Volume 1: Hansen, Geschichte der Stadt Narva (1858)
This initial release includes the "Einleitung" (introduction) section of the 1858 book Geschichte der Stadt Narva by H. J. Hansen, a Danish-German ethnographer. The dataset includes:
- 87 aligned records
- OCR text (Fraktur)
- Normalized German
- 7 parallel translations
Format
Each .jsonl and .csv entry includes the following fields:
original_frakturnormalized_germantranslation_entranslation_rutranslation_ettranslation_svtranslation_fitranslation_da
License
CC BY-SA 4.0 You are free to use, adapt, and redistribute with proper attribution.
Author & Contact
This dataset is created by k8mpass, an independent open-source initiative based in Estonia. Feel free to cite, fork, or contact for collaborations.
📥 Download instructions
- Use
dataset.csvto explore or visualize the corpus. - To use the full dataset for machine learning training:
- Download `data.zip`
- Unzip it to extract the file
data.jsonl - Load
data.jsonlusing your preferred JSONL reader or training pipeline - File is encoded in UTF-8 and follows JSON Lines format (one JSON object per line).
