bavarian-nlp/bavarian-books-ocred-v0.1
📚 🥨 OCR'ed Bavarian Books Due to the lack of high quality resources for Bavarian, I've started this dataset repo for OCR'ing Bavarian books. The dataset is based on the Bavarian Books dataset. 🧮 OCR The current form of this dataset uses the awesome Tesseract library for OCR'ing the Bavarian books. We use the following Fraktur model: wget https://github.com/tesseract-ocr/tessdata/raw/refs/heads/main/script/Fraktur.traineddata 📄 Dataset Format… See the full description on the dataset page: https://huggingface.co/datasets/bavarian-nlp/bavarian-books-ocred-v0.1.
07
feat: add initial version of dataset
docs: use more emojis
docs: minor typo fix
docs: add initial version
initial commit
