CoolFace
Datasetpublic

Srijan-Upadhyay/arxiv-page-ocr-bench

arXiv Page OCR Benchmark Corpus Dataset Summary arXiv Page OCR Benchmark Corpus is an evaluation benchmark dataset containing high-resolution rendered document page images from scientific arXiv papers (covering STEM disciplines including physics, computer science, mathematics, and quantitative finance) paired with ground-truth Markdown / LaTeX text representations. This dataset is designed to evaluate document parsing, optical character recognition (OCR), and… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Upadhyay/arxiv-page-ocr-bench.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes77downloads

Srijan-Upadhyay/arxiv-page-ocr-bench · main · files are served by the source, never re-hosted here