llm-jp/Synth-JDoc
Synth-JDoc Paper | Code Synth-JDoc is a dataset of synthetic Japanese document images generated using HTML/CSS. We generate embedded images, captions, and titles from prepared text, and use these elements to synthesize document images featuring diverse multi-column layouts in both vertical and horizontal writing. Because the document images are synthesized directly from text, this dataset is completely free from OCR errors. Dataset details id Image ID image… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/Synth-JDoc.
This repository belongs to llm-jp on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
