CoolFace
Datasetpublic

thekamilya/kazakh-printed-dataset

Kazakh Printed Dataset for OCR task Data Lineage This dataset was synthetically generated using issai/kazparc as the base. Since Kazakh OCR data is scarce, I developed a pipeline to transform digital Kazakh text into a printed-style dataset. Generation Process Source: Text samples were extracted from issai/kazparc. Augmentation & Stylization: Random Background Color: Simulates different lighting conditions by… See the full description on the dataset page: https://huggingface.co/datasets/thekamilya/kazakh-printed-dataset.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes23downloads

thekamilya/kazakh-printed-dataset · main · files are served by the source, never re-hosted here