boydcheung/TAL_OCR_Composed_37K
Dataset Information TAL_OCR_Composed_37K is a composed set of TAL OCR datasets. The original images and labels are downloaded from TAL website (https://ai.100tal.com/dataset), includes mainly K12 Handwritten Chinese texts, English texts, Math formulas. Printed K12 materials. Images are tiled randomly to have a more compact view. There are total 32K images and text pairs after processing: TAL_OCR_CHN/composed (645 images) TAL_OCR_ENG/composed (399 images)… See the full description on the dataset page: https://huggingface.co/datasets/boydcheung/TAL_OCR_Composed_37K.
025
