CoolFace
Datasetpublic

HumynLabs/Arabic_Documents_Dataset_PDF

Arabic Documents Dataset (PDF) This dataset contains a collection of Arabic-language documents in PDF format. The corpus includes books, articles, reports, and educational materials written in Modern Standard Arabic and regional variants. It is curated to support AI research in document understanding, Arabic OCR, and text extraction from complex layouts. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Arabic_Documents_Dataset_PDF.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
0likes1.2kdownloads
4 commits on main
078cf8811mo ago

Update README.md

KAI
b3d5e3211mo ago

Upload 127 files

KAI
35ec66d11mo ago

Update README.md

KAI
764b27811mo ago

initial commit

KAI