CoolFace
Datasetpublic

HumynLabs/Chinese_Documents_Dataset_PDF

Chinese Documents Dataset (PDF) This dataset consists of a curated collection of Chinese-language documents in PDF format. It includes textbooks, research papers, articles, public-domain books, and official documents written in Simplified and Traditional Chinese. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Chinese_Documents_Dataset_PDF.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
0likes647downloads
settings

This repository belongs to HumynLabs on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameChinese_Documents_Dataset_PDF
visibilitypublic
licencecc-by-4.0
gatedno
ownerHumynLabs
Account settings
HumynLabs/Chinese_Documents_Dataset_PDF · CoolFace