CoolFace
Datasetpublic

ramSeraph/indic_wikisource

Dataset contains page image urls and the corresponding annotations from wikisource. Also has information whether the page has been validated/proofread. I expect this dataset to be useful for creating OCR models for printed text in Indic languages. Languages Covered: as - Assamese bn - Bengali gu - Gujarati hi - Hindi kn - Kannada ml - Malayalam mr - Marathi or - Odiya pa - Punjabi sa - Sanskrit ta - Tamil te - Telugu

sourceHugging Facecc-by-sa-4.0updated 2y agoView on Hugging Face
1likes126downloads
6 commits on main
c4212542y ago

Update Readme

ramSeraph
f35e8642y ago

Uploading the remaining files

ramSeraph
316f3ee2y ago

Delete te.pq

ramSeraph
1ad6a9a2y ago

upload te.parquet

ramSeraph
8fee65a2y ago

Upload te.pq

ramSeraph
a5fe3142y ago

initial commit

ramSeraph