CoolFace
Datasetpublic

Sigurdur/icelandic-ocr-benchmark

Dataset Card for Icelandic OCR Benchmark Dataset Details Dataset Description Icelandic OCR Benchmark is a ground-truth dataset for evaluating OCR accuracy on Icelandic-language documents. It consists of manually transcribed page images with matching layout annotations (text regions, line polygons, baselines) in both ALTO and PAGE XML. Curated by: Sigurdur Haukur Birgisson Language(s): Icelandic (is) License: CC BY-SA 4.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/icelandic-ocr-benchmark.

sourceHugging Facecc-by-sa-4.0updated 10d agoView on Hugging Face
1likes87downloads
4 commits on main
441026a10d ago

Add dataset card

Sigurdur
f6cb3fb10d ago

Set license to cc-by-sa-4.0

Sigurdur
8ec06b610d ago

Add train split from eScriptorium

Sigurdur
101c3c210d ago

initial commit

Sigurdur