datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chinese-traditional-culturetraditional-chinese-ocr-synthetic
Traditional Chinese OCR Synthetic Dataset
A large-scale synthetic dataset containing 4.1 million image-text pairs specifically designed for Traditional Chinese historical document recognition.
Dataset Overview
Existing large-scale Traditional Chinese OCR datasets (e.g., TCSynth) are primarily designed for scene text recognition, characterized by:
Horizontal layouts
Short text sequences (2-5 characters on average)
Modern commonly-used characters
These characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ZihCiLin/traditional-chinese-ocr-synthetic.traditional-chinese-historical-ocr-lo-chia-luen
Traditional Chinese Historical OCR Dataset
(Lo Chia-Lun Manuscripts)
This dataset consists of manually annotated OCR text crops derived from the Lo Chia-Lun Manuscript Collection (羅家倫文稿), hosted by the National Chengchi University Library.
The dataset is designed to support research on Traditional Chinese OCR, particularly for historical documents characterized by vertical layouts, handwritten or semi-printed glyphs, and long-form text lines.
Due to archival and… See the full description on the dataset page: https://huggingface.co/datasets/ZihCiLin/traditional-chinese-historical-ocr-lo-chia-luen.content.chinese_traditional
