CoolFace
11 results

transnormer

textplus-bbaw /dtak-transnormer-basic-v1 Dataset Card for DTAK-transnormer-basic (v1.0) Dataset Details Dataset Description DTAK-transnormer-basic is a modified subset of the DTA-Kernkorpus (Deutsches Textarchiv, German Text Archive Core Corpus). It is a parallel corpus of German texts from the period between 1600 to 1899, that aligns sentences in historical spelling with their normalizations. A normalization is a modified version of the original text that is adapted to modern spelling conventions.… See the full description on the dataset page: https://huggingface.co/datasets/textplus-bbaw/dtak-transnormer-basic-v1.tabular1M<n<10M0 likes96 downloads1y agoHugging Facetextplus-bbaw /dtak-transnormer-full-v1 Dataset Card for DTAK-transnormer-full (v1.0) Dataset Details Dataset Description DTAK-transnormer-full is a modified subset of the DTA-Kernkorpus (Deutsches Textarchiv, German Text Archive Core Corpus). It is a parallel corpus of German texts from the period between 1600 to 1899, that aligns sentences in historical spelling with their normalizations. A normalization is a modified version of the original text that is adapted to modern spelling conventions. This… See the full description on the dataset page: https://huggingface.co/datasets/textplus-bbaw/dtak-transnormer-full-v1.0 likes55 downloads2y agoHugging Faceybracke /lexicon-dtak-transnormer-v1 Lexicon-DTAK-transnormer (v1.0) This dataset is derived from dtak-transnormer-full-v1, a parallel corpus of German texts from the period between 1600 to 1899, that aligns sentences in historical spelling with their normalizations. This dataset is a lexicon of ngram alignments between original and normalized ngrams observed in dtak-transnormer-full-v1 and their frequency. The ngram alignments in the lexicon are drawn from the sentence-level ngram alignments in… See the full description on the dataset page: https://huggingface.co/datasets/ybracke/lexicon-dtak-transnormer-v1.tabular1M<n<10M0 likes13 downloads2y agoHugging Face