Nayana-cognitivelab/NayanaDocs-Global-45k-webdataset
Nayana-DocOCR Global Annotated Dataset Dataset Description This is a large-scale multilingual document OCR dataset containing approximately 400GB of images with comprehensive annotations across multiple global languages and English. The dataset is stored in WebDataset format using TAR archives for efficient streaming and processing. Available Language Subsets Arabic (ar): Available German (de): Available Russian (ru) : Available Spanish (es):… See the full description on the dataset page: https://huggingface.co/datasets/Nayana-cognitivelab/NayanaDocs-Global-45k-webdataset.
14
No card is published for this repository, or it could not be fetched from Hugging Face right now.
