bavarian-nlp/bavarian-books-ocred-v0.1
📚 🥨 OCR'ed Bavarian Books Due to the lack of high quality resources for Bavarian, I've started this dataset repo for OCR'ing Bavarian books. The dataset is based on the Bavarian Books dataset. 🧮 OCR The current form of this dataset uses the awesome Tesseract library for OCR'ing the Bavarian books. We use the following Fraktur model: wget https://github.com/tesseract-ocr/tessdata/raw/refs/heads/main/script/Fraktur.traineddata 📄 Dataset Format… See the full description on the dataset page: https://huggingface.co/datasets/bavarian-nlp/bavarian-books-ocred-v0.1.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face