datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dtak-transnormer-basic-v1
Dataset Card for DTAK-transnormer-basic (v1.0)
Dataset Details
Dataset Description
DTAK-transnormer-basic is a modified subset of the DTA-Kernkorpus (Deutsches Textarchiv, German Text Archive Core Corpus).
It is a parallel corpus of German texts from the period between 1600 to 1899, that aligns sentences in historical spelling with their normalizations.
A normalization is a modified version of the original text that is adapted to modern spelling conventions.… See the full description on the dataset page: https://huggingface.co/datasets/textplus-bbaw/dtak-transnormer-basic-v1.dtak-transnormer-full-v1
Dataset Card for DTAK-transnormer-full (v1.0)
Dataset Details
Dataset Description
DTAK-transnormer-full is a modified subset of the DTA-Kernkorpus (Deutsches Textarchiv, German Text Archive Core Corpus).
It is a parallel corpus of German texts from the period between 1600 to 1899, that aligns sentences in historical spelling with their normalizations.
A normalization is a modified version of the original text that is adapted to modern spelling conventions.
This… See the full description on the dataset page: https://huggingface.co/datasets/textplus-bbaw/dtak-transnormer-full-v1.lexicon-dtak-transnormer-v1
Lexicon-DTAK-transnormer (v1.0)
This dataset is derived from dtak-transnormer-full-v1, a parallel corpus of German texts from the period between 1600 to 1899, that aligns sentences in historical spelling with their normalizations.
This dataset is a lexicon of ngram alignments between original and normalized ngrams observed in dtak-transnormer-full-v1 and their frequency. The ngram alignments in the lexicon are drawn from the sentence-level ngram alignments in… See the full description on the dataset page: https://huggingface.co/datasets/ybracke/lexicon-dtak-transnormer-v1.
