CoolFace
Datasetpublic

samaritan-ai/samaritan_hebrew_LightOnOcr

Samaritan Hebrew OCR Dataset Dataset Summary The Samaritan Hebrew OCR Dataset is a specialized dataset for fine-tuning OCR models on Samaritan Hebrew manuscripts. This dataset contains 46,860 annotated samples extracted from 1,374 manuscript pages, converted from PAGE-XML format to the LightOnOCR-2 training format. The dataset includes three types of samples: Line-level samples: Individual textlines cropped using precise polygon masks (40,219 samples)… See the full description on the dataset page: https://huggingface.co/datasets/samaritan-ai/samaritan_hebrew_LightOnOcr.

sourceHugging Facecc-by-4.0updated 8mo agoView on Hugging Face
0likes21downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
samaritan-ai/samaritan_hebrew_LightOnOcr · CoolFace