CoolFace
Datasetpublic

thekamilya/kazakh-printed-dataset

Kazakh Printed Dataset for OCR task Data Lineage This dataset was synthetically generated using issai/kazparc as the base. Since Kazakh OCR data is scarce, I developed a pipeline to transform digital Kazakh text into a printed-style dataset. Generation Process Source: Text samples were extracted from issai/kazparc. Augmentation & Stylization: Random Background Color: Simulates different lighting conditions by… See the full description on the dataset page: https://huggingface.co/datasets/thekamilya/kazakh-printed-dataset.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes23downloads
settings

This repository belongs to thekamilya on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namekazakh-printed-dataset
visibilitypublic
licenceapache-2.0
gatedno
ownerthekamilya
Account settings
thekamilya/kazakh-printed-dataset · CoolFace