CoolFace
Datasetpublic

ReadingTimeMachine/rtm-sgt-ocr-v1

Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic Data from the ar𝜒iv for OCR Post Correction of Historic Scientific Articles". Synthetic ground truth (SGT) sentences have been mined from the ar𝜒iv Bulk Downloads source documents, and Optical Character Recognition (OCR) sentences have been generated with the Tesseract OCR engine on the PDF pages generated from compiled source documents.… See the full description on the dataset page: https://huggingface.co/datasets/ReadingTimeMachine/rtm-sgt-ocr-v1.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
4likes679downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

ReadingTimeMachine/rtm-sgt-ocr-v1 · main · files are served by the source, never re-hosted here