CoolFace
Datasetpublic

bavarian-nlp/bavarian-books-ocred-v0.1

📚 🥨 OCR'ed Bavarian Books Due to the lack of high quality resources for Bavarian, I've started this dataset repo for OCR'ing Bavarian books. The dataset is based on the Bavarian Books dataset. 🧮 OCR The current form of this dataset uses the awesome Tesseract library for OCR'ing the Bavarian books. We use the following Fraktur model: wget https://github.com/tesseract-ocr/tessdata/raw/refs/heads/main/script/Fraktur.traineddata 📄 Dataset Format… See the full description on the dataset page: https://huggingface.co/datasets/bavarian-nlp/bavarian-books-ocred-v0.1.

sourceHugging Faceotherupdated 1y agoView on Hugging Face
0likes7downloads
settings

This repository belongs to bavarian-nlp on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namebavarian-books-ocred-v0.1
visibilitypublic
licenceother
gatedno
ownerbavarian-nlp
Account settings
bavarian-nlp/bavarian-books-ocred-v0.1 · CoolFace