ramSeraph/indic_wikisource
Dataset contains page image urls and the corresponding annotations from wikisource. Also has information whether the page has been validated/proofread. I expect this dataset to be useful for creating OCR models for printed text in Indic languages. Languages Covered: as - Assamese bn - Bengali gu - Gujarati hi - Hindi kn - Kannada ml - Malayalam mr - Marathi or - Odiya pa - Punjabi sa - Sanskrit ta - Tamil te - Telugu
1126
Dataset contains page image urls and the corresponding annotations from wikisource. Also has information whether the page has been validated/proofread.
I expect this dataset to be useful for creating OCR models for printed text in Indic languages.
Languages Covered:
- as - Assamese
- bn - Bengali
- gu - Gujarati
- hi - Hindi
- kn - Kannada
- ml - Malayalam
- mr - Marathi
- or - Odiya
- pa - Punjabi
- sa - Sanskrit
- ta - Tamil
- te - Telugu
