allenai/olmOCR-bench
olmOCR-bench olmOCR-bench is a dataset of 1,403 PDF files, plus 7,010 unit test cases that capture properties of the output that a good OCR system should have. This benchmark evaluates the ability of OCR systems to accurately convert PDF documents to markdown format while preserving critical textual and structural information. Quick links: 📃 Paper 🛠️ Code 🎮 Demo Table 1. Distribution of Test Classes by Document Source Document Source Text Present Text… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-bench.
Update eval.yaml
Update eval.yaml
Create eval.yaml
Update README.md
Updated links for 11_*.pdf in long tiny text (#2)
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Final removal of pii
removed PII content
Removing one document that overlapped with olmocr-mix training set
updated multi_column after failed tests
updated pdf's to high quality
Some small fixes
fixed (,.) pattern's in test cases
Fix some headers footers issues
Supper
Fix duplicate id
Marking all absent tests as not case sensitive by default
Adding two baseline check overrides
removed rejected cases
updated multi_column tests
removed duplicates
removed duplicates
cleaned multi-column
Rejected table tests removed
fixed old_scans
Merge branch 'main' of https://huggingface.co/datasets/allenai/olmOCR-bench into main
One tiny fix
Fixed 91 table tests, and 14 order tests
Merge branch 'main' of https://huggingface.co/datasets/allenai/olmOCR-bench
Cleanup for math and table tests
cleaned headers_footers data from o3
Small cleanup review meeting
updated old_scans_math data
cleaned old_scans data
Cleaned up long texts
Merge branch 'main' of https://huggingface.co/datasets/allenai/olmOCR-bench
Verified more
added a missing pdf of long_tiny_text
A few more corrections
Some hyphen corrections
Cleaning some longer tests
added tiny_long_text test cases
Cleaning up arxiv maths not to include ending punctuations
