CoolFace
Datasetpublic

avishadilhara/sinhala-ocr-lk-acts-1010

🇱🇰 Sinhala OCR - Sri Lankan Acts Dataset Dataset Description This dataset contains 1,010 scanned document images of Sri Lankan legal acts (1980s-2010s) in Sinhala language with ground truth text annotations for Optical Character Recognition (OCR) training and evaluation. Key Features ✅ High-quality scanned document images ✅ Professionally corrected ground truth text ✅ Year-wise metadata for temporal analysis ✅ Pre-split into train/eval/test sets… See the full description on the dataset page: https://huggingface.co/datasets/avishadilhara/sinhala-ocr-lk-acts-1010.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
1likes135downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
avishadilhara/sinhala-ocr-lk-acts-1010 · CoolFace