CoolFace
Datasetpublic

Nishant2414/OCR-Synthetic-Multilingual-v1

OCR-Synthetic-Multilingual-v1 Overview Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset page: https://huggingface.co/datasets/Nishant2414/OCR-Synthetic-Multilingual-v1.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes14kdownloads
settings

This repository belongs to Nishant2414 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameOCR-Synthetic-Multilingual-v1
visibilitypublic
licencecc-by-4.0
gatedno
ownerNishant2414
Account settings
Nishant2414/OCR-Synthetic-Multilingual-v1 · CoolFace