CoolFace
Datasetpublicgated

Nayana-cognitivelab/NayanaDocs-Global-45k-webdataset

Nayana-DocOCR Global Annotated Dataset Dataset Description This is a large-scale multilingual document OCR dataset containing approximately 400GB of images with comprehensive annotations across multiple global languages and English. The dataset is stored in WebDataset format using TAR archives for efficient streaming and processing. Available Language Subsets Arabic (ar): Available German (de): Available Russian (ru) : Available Spanish (es):… See the full description on the dataset page: https://huggingface.co/datasets/Nayana-cognitivelab/NayanaDocs-Global-45k-webdataset.

sourceHugging Faceupdated 1y agoView on Hugging Face
1likes4downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Nayana-cognitivelab/NayanaDocs-Global-45k-webdataset · CoolFace