CoolFace
Datasetpublic

Sigurdur/icelandic-ocr-benchmark

Dataset Card for Icelandic OCR Benchmark Dataset Details Dataset Description Icelandic OCR Benchmark is a ground-truth dataset for evaluating OCR accuracy on Icelandic-language documents. It consists of manually transcribed page images with matching layout annotations (text regions, line polygons, baselines) in both ALTO and PAGE XML. Curated by: Sigurdur Haukur Birgisson Language(s): Icelandic (is) License: CC BY-SA 4.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/icelandic-ocr-benchmark.

sourceHugging Facecc-by-sa-4.0updated 11d agoView on Hugging Face
1likes87downloads
../
filetrain-00000-of-00001.parquet45.7 MBdownload

Sigurdur/icelandic-ocr-benchmark · main · files are served by the source, never re-hosted here