ramSeraph/indic_wikisource
Dataset contains page image urls and the corresponding annotations from wikisource. Also has information whether the page has been validated/proofread. I expect this dataset to be useful for creating OCR models for printed text in Indic languages. Languages Covered: as - Assamese bn - Bengali gu - Gujarati hi - Hindi kn - Kannada ml - Malayalam mr - Marathi or - Odiya pa - Punjabi sa - Sanskrit ta - Tamil te - Telugu
1126
Update Readme
Uploading the remaining files
Delete te.pq
upload te.parquet
Upload te.pq
initial commit
