CoolFace
Datasetpublic

ramSeraph/indic_wikisource

Dataset contains page image urls and the corresponding annotations from wikisource. Also has information whether the page has been validated/proofread. I expect this dataset to be useful for creating OCR models for printed text in Indic languages. Languages Covered: as - Assamese bn - Bengali gu - Gujarati hi - Hindi kn - Kannada ml - Malayalam mr - Marathi or - Odiya pa - Punjabi sa - Sanskrit ta - Tamil te - Telugu

sourceHugging Facecc-by-sa-4.0updated 2y agoView on Hugging Face
1likes126downloads
Dataset Card

Dataset contains page image urls and the corresponding annotations from wikisource. Also has information whether the page has been validated/proofread.

I expect this dataset to be useful for creating OCR models for printed text in Indic languages.

Languages Covered:

  • —as - Assamese
  • —bn - Bengali
  • —gu - Gujarati
  • —hi - Hindi
  • —kn - Kannada
  • —ml - Malayalam
  • —mr - Marathi
  • —or - Odiya
  • —pa - Punjabi
  • —sa - Sanskrit
  • —ta - Tamil
  • —te - Telugu