CoolFace
Datasetpublic

ReadingTimeMachine/rtm-sgt-ocr-v1

Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic Data from the ar𝜒iv for OCR Post Correction of Historic Scientific Articles". Synthetic ground truth (SGT) sentences have been mined from the ar𝜒iv Bulk Downloads source documents, and Optical Character Recognition (OCR) sentences have been generated with the Tesseract OCR engine on the PDF pages generated from compiled source documents.… See the full description on the dataset page: https://huggingface.co/datasets/ReadingTimeMachine/rtm-sgt-ocr-v1.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
4likes679downloads
7 commits on main
e0c90f41y ago

Add OCR tag (#2)

jnaiman, librarian-bot
f87fbcc3y ago

Update README.md

jnaiman
aff40a13y ago

Update README.md

jnaiman
702c4253y ago

Upload 153 files

jnaiman
a0036c93y ago

Update README.md

jnaiman
f6a267f3y ago

Update README.md

jnaiman
1dd5fef3y ago

initial commit

jnaiman