CoolFace
Datasetpublic

Srijan-Upadhyay/arxiv-page-ocr-bench

arXiv Page OCR Benchmark Corpus Dataset Summary arXiv Page OCR Benchmark Corpus is an evaluation benchmark dataset containing high-resolution rendered document page images from scientific arXiv papers (covering STEM disciplines including physics, computer science, mathematics, and quantitative finance) paired with ground-truth Markdown / LaTeX text representations. This dataset is designed to evaluate document parsing, optical character recognition (OCR), and… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Upadhyay/arxiv-page-ocr-bench.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes77downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Srijan-Upadhyay/arxiv-page-ocr-bench · CoolFace