CoolFace
Datasetpublic

bavarian-nlp/bavarian-books-ocred-v0.1

📚 🥨 OCR'ed Bavarian Books Due to the lack of high quality resources for Bavarian, I've started this dataset repo for OCR'ing Bavarian books. The dataset is based on the Bavarian Books dataset. 🧮 OCR The current form of this dataset uses the awesome Tesseract library for OCR'ing the Bavarian books. We use the following Fraktur model: wget https://github.com/tesseract-ocr/tessdata/raw/refs/heads/main/script/Fraktur.traineddata 📄 Dataset Format… See the full description on the dataset page: https://huggingface.co/datasets/bavarian-nlp/bavarian-books-ocred-v0.1.

sourceHugging Faceotherupdated 1y agoView on Hugging Face
0likes7downloads
5 commits on main
7a6004a1y ago

feat: add initial version of dataset

stefan-it
c8769271y ago

docs: use more emojis

stefan-it
3ab9d141y ago

docs: minor typo fix

stefan-it
eb6cd321y ago

docs: add initial version

stefan-it
bbe6a141y ago

initial commit

stefan-it