CoolFace
Datasetpublic

llm-jp/Synth-JDoc

Synth-JDoc Paper | Code Synth-JDoc is a dataset of synthetic Japanese document images generated using HTML/CSS. We generate embedded images, captions, and titles from prepared text, and use these elements to synthesize document images featuring diverse multi-column layouts in both vertical and horizontal writing. Because the document images are synthesized directly from text, this dataset is completely free from OCR errors. Dataset details id Image ID image… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/Synth-JDoc.

sourceHugging Facecc-by-4.0updated 26d agoView on Hugging Face
0likes430downloads
settings

This repository belongs to llm-jp on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameSynth-JDoc
visibilitypublic
licencecc-by-4.0
gatedno
ownerllm-jp
Account settings
llm-jp/Synth-JDoc · CoolFace