CoolFace
Datasetpublic

HumynLabs/Japanese_Documents_Dataset_PDF

Japanese Documents Dataset (PDF) This dataset contains a curated collection of Japanese-language documents in PDF format. The corpus includes textbooks, research papers, news articles, public-domain books, and government publications written in Japanese. It is intended to support AI research in OCR, document understanding, and multilingual text recognition. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Japanese_Documents_Dataset_PDF.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
2likes665downloads
settings

This repository belongs to HumynLabs on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameJapanese_Documents_Dataset_PDF
visibilitypublic
licencecc-by-4.0
gatedno
ownerHumynLabs
Account settings
HumynLabs/Japanese_Documents_Dataset_PDF · CoolFace