CoolFace
Datasetpublic

taresco/KarantaOCR-Bench

KarantaOCR - Bench KarantaOCR-Bench is a unit-test–style evaluation dataset, similar to olmOCR-bench. It consists of 70 PDF documents and 300 test cases spanning multiple document types. All the tests were manually verified by us. KarantaOCR-Bench is designed specifically to evaluate document text extraction for Documents with diacritics and special characters, covering a diverse range of document formats and languages commonly under-represented in existing OCR benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/taresco/KarantaOCR-Bench.

sourceHugging Faceodc-byupdated 8mo agoView on Hugging Face
0likes396downloads
11 commits on main
4abbcbe8mo ago

Update README.md

ToluClassics
7a721c28mo ago

add new category

ToluClassics
7924fbf8mo ago

add new category

ToluClassics
8305a6f9mo ago

Update README.md

ToluClassics
ccca0689mo ago

Update README.md

ToluClassics
37d98649mo ago

Update README.md

ToluClassics
6ba51aa9mo ago

Update README.md

ToluClassics
4085ce49mo ago

Upload 1896 files

ToluClassics
331b4e39mo ago

Delete bench-data

ToluClassics
be0e0399mo ago

Upload 2119 files

ToluClassics
8f2a9279mo ago

initial commit

ToluClassics