datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sinhala_OCR_Dataset_SyntheticThis is a synthetically generated dataset you can generate your own with the help of below codes in this github repo: https://github.com/suchinthana00/Synthetic_OCR_Dataset_Generator
sinhala-ocr-lk-acts-1010
🇱🇰 Sinhala OCR - Sri Lankan Acts Dataset
Dataset Description
This dataset contains 1,010 scanned document images of Sri Lankan legal acts (1980s-2010s) in Sinhala language with ground truth text annotations for Optical Character Recognition (OCR) training and evaluation.
Key Features
✅ High-quality scanned document images
✅ Professionally corrected ground truth text
✅ Year-wise metadata for temporal analysis
✅ Pre-split into train/eval/test sets… See the full description on the dataset page: https://huggingface.co/datasets/avishadilhara/sinhala-ocr-lk-acts-1010.sinhala_synthetic_ocr_news_largeCreated from https://www.kaggle.com/code/ransakaravihara/sinhala-ocr-image-creation.
To cite the dataset
@misc{ransaka_r._2026,
author = { Ransaka R. },
title = { sinhala_synthetic_ocr_news_large (Revision bc52307) },
year = 2026,
url = { https://huggingface.co/datasets/edifier99/sinhala_synthetic_ocr_news_large },
doi = { 10.57967/hf/9748 },
publisher = { Hugging Face }
}
sinhala-handwritten-ocr-623sinhala_synthetic_ocr-largeIf you use this data in publications, please cite it as follows:
@misc {ransaka_ravihara_2024,
author = { {Ransaka Ravihara} },
title = { sinhala_synthetic_ocr-large (Revision f3cac3b) },
year = 2024,
url = { https://huggingface.co/datasets/Ransaka/sinhala_synthetic_ocr-large },
doi = { 10.57967/hf/1809 },
publisher = { Hugging Face }
}
sinhala-ocrsinhala_synthetic_ocrsinhala_synthetic_ocr-newssinhala-ocr-synthetic-new
