CoolFace
8 results

fraktur

Riksarkivet /swedish_fraktur Swedish Fraktur This is a dataset for swedish blackletter from the 19th century. The transcriptions were made by Språkbanken and converted into a text-line dataset by the Swedish National Archives Dataset Details Uses Direct Use Train textline-based OCR models for swedish 19th century blackletter Dataset Structure { "image": Image(), "text": str } Dataset Creation Source Data Original dataset from Språkbanken -… See the full description on the dataset page: https://huggingface.co/datasets/Riksarkivet/swedish_fraktur.imageimage-to-text10K<n<100K2 likes163 downloads2y agoHugging Faceimpresso-project /frakturline-testset Fraktur/Other Text-Line — Test Set A balanced, held-out evaluation set of 2 000 scanned text-line images (1 000 per class) for the binary task of distinguishing Fraktur (blackletter / Gothic script) from other script (primarily Antiqua / Latin / Roman). Developed for the Impresso digital humanities project. Dataset Details Property Value Task Binary image classification Classes fraktur, other Images per class 1 000 Total images 2 000 Image format WebP… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/frakturline-testset.imageimage-classification1K<n<10K1 likes145 downloads6mo agoHugging Faceimpresso-project /frakturline-dataset Fraktur/Other Text-Line Classifier A binary CNN classifier that determines whether a scanned text-line image is set in Fraktur (blackletter / Gothic script) or Other (primarily Latin / Roman / Antiqua script). Developed for the Impresso digital humanities project, which processes millions of historical newspaper pages in German, French, Luxembourgish, and other European languages. Model Details Property Value Architecture BinaryClassificationCNN — 3-layer… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/frakturline-dataset.image10K<n<100K0 likes33 downloads6mo agoHugging Facek8mpass-fraktur-series /fraktur-baltic-corpus Fraktur Baltic Corpus Fraktur Baltic Corpus is a multilingual dataset based on historical German-language books printed in the Russian Empire during the 18th–19th centuries, primarily in Fraktur typeface. Each entry in the dataset contains: Raw OCR text from historical Fraktur sources Normalized German version Translations into seven languages: English, Russian, Estonian, Swedish, Finnish, Danish, and Modern German Volume 1: Hansen, Geschichte der Stadt Narva (1858)… See the full description on the dataset page: https://huggingface.co/datasets/k8mpass-fraktur-series/fraktur-baltic-corpus.texttranslationn<1K1 likes24 downloads1y agoHugging FaceV4ldeLund /Swedish-fraktur-1626-1816 Swedish Fraktur 1626-1816 Processed page-level OCR dataset built from Svensk fraktur 1626-1816 with one scanned page image paired to one manual transcription file. Data citation Information Språkbanken Text (2021). Swedish fraktur 1626-1816 (updated: 2021-11-26). [Data set]. Språkbanken Text. https://doi.org/10.23695/5sme-7437 Dataset Structure This parquet file contains the following columns: image (image): page image bytes text (string): full page… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/Swedish-fraktur-1626-1816.imagen<1K0 likes8 downloads7mo agoHugging Face