datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kat57-ocr-bench-500-results
Kat57 OCR benchmark — CER/WER
Strict reference-based evaluation of 16 OCR models on a deterministic 500-card sample from Lund University Library's Kat57 catalogue-card collection. This result set contains only Character Error Rate (CER) and Word Error Rate (WER); it does not contain VLM judging or ELO ratings.
The sample was drawn with seed 57 from tadad/kat57-ground-truth and is published as tadad/kat57-ground-truth-500. The OCR outputs are retained in… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-500-results.kat57-ocr-bench-results
Kat57 OCR smoke benchmark results
Exact ground-truth scoring for a 50-card Tesseract integration run over tadad/kat57-ground-truth-smoke. OCR outputs are published in the tesseract config of tadad/kat57-ocr-bench.
Model
CER
WER
Evaluated
Empty outputs
Error sentinels
Skipped references
Tesseract 5
0.4656
0.8605
50
1
0
0
The corpus totals are 4,629 character edits over 9,941 reference characters and 1,221 word edits over 1,419 reference words. Scoring used… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-results.ocr-bench-britannica-results-qwen35
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
zai-org/GLM-OCR
0.9B
1716
1673–1769
182
60
2
75%
2
lightonai/LightOnOCR-2-1B
1B
1697
1655–1749
158
61
1
72%
3
numind/NuExtract3
4B
1649
1604–1708
162
82
0
66%
4
FireRedTeam/FireRed-OCR
2.1B
1506
1469–1548
115
127
2
47%
5… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-britannica-results-qwen35.bpl-ocr-bench-results
OCR Bench Results: bpl-ocr-bench
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
ELO
95% CI
Wins
Losses
Ties
Win%
1
lightonai/LightOnOCR-2-1B
1559
1497–1630
39
25
0
61%
2
zai-org/GLM-OCR
1535
1471–1591
48
35
1
57%
3
rednote-hilab/dots.ocr
1453
1385–1515
26
37
0
41%
4
deepseek-ai/DeepSeek-OCR
1452
1388–1514
33
49
1
40%
Details
Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/bpl-ocr-bench-results.ocr-bench-judge-eval-27b
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
ELO
95% CI
Wins
Losses
Ties
Win%
1
lightonai/LightOnOCR-2-1B
1675
1571–1836
26
9
1
72%
2
FireRedTeam/FireRed-OCR
1612
1518–1767
25
13
1
64%
3
zai-org/GLM-OCR
1594
1480–1739
24
14
1
62%
4
deepseek-ai/DeepSeek-OCR
1437
1332–1546
15
23
1
38%
5
rednote-hilab/dots.ocr
1182… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-judge-eval-27b.ocr-bench-britannica-results
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
rednote-hilab/dots.mocr
3B
1745
1714–1782
436
141
4
75%
2
lightonai/LightOnOCR-2-1B
1B
1741
1709–1779
426
141
4
75%
3
zai-org/GLM-OCR
0.9B
1738
1707–1773
469
157
2
75%
4
allenai/olmOCR-2-7B-1025-FP8
1719
1688–1753
454
167
4
73%… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-britannica-results.ocr-bench-ufo-results
OCR Bench Results: ocr-bench-ufo
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
ELO
95% CI
Wins
Losses
Ties
Win%
1
deepseek-ai/DeepSeek-OCR
1691
1620–1801
46
12
0
79%
2
lightonai/LightOnOCR-2-1B
1570
1492–1661
36
23
0
61%
3
rednote-hilab/dots.ocr
1432
1339–1512
22
35
1
38%
4
zai-org/GLM-OCR
1307
1209–1374
11
45
1
19%
Details
Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-ufo-results.ocr-benchmark-results
OCR Bench Results: ocr-benchmark-combined
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
clearocr.com/clearocr-api
1837
1801–1882
488
56
1
90%
2
deepseek-ai/DeepSeek-OCR
4B
1639
1611–1671
380
164
1
70%
3
lightonai/LightOnOCR-2-1B
1B
1448
1421–1477
242
302
0
44%
4
rednote-hilab/dots.ocr
1.7B
1384
1353–1413
193
351
0
35%… See the full description on the dataset page: https://huggingface.co/datasets/j4xfu2mm/ocr-benchmark-results.ocr-bench-judge-eval-122b
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
ELO
95% CI
Wins
Losses
Ties
Win%
1
lightonai/LightOnOCR-2-1B
1833
1725–2478
31
4
1
86%
2
zai-org/GLM-OCR
1610
1476–2239
24
14
1
62%
3
FireRedTeam/FireRed-OCR
1571
1465–2167
22
16
1
56%
4
deepseek-ai/DeepSeek-OCR
1431
1292–2034
15
23
1
38%
5
rednote-hilab/dots.ocr
1054… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-judge-eval-122b.ocr-bench-moh-results
OCR Bench Results: ocr-bench-moh
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
lightonai/LightOnOCR-2-1B
1B
1618
1592–1643
410
210
21
64%
2
deepseek-ai/DeepSeek-OCR-2
3.4B
1579
1555–1604
385
250
14
59%
3
baidu/Qianfan-OCR
4.7B
1577
1550–1604
377
250
11
59%
4
numind/NuExtract3
4B
1572
1549–1598
370
249
30
57%
5… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-moh-results.ocr-bench-rubenstein-judgeocr-bench-ufo-judge-30bocr-bench-britannica-nuextract3-sweep
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
glm-ocr
1556
1513–1599
116
78
0
60%
2
lighton-ocr-2
1538
1496–1580
99
76
0
57%
3
nuextract3
1532
1495–1573
104
82
7
54%
4
nuextract3-t0rep
1491
1451–1531
91
97
6
47%
5
nuextract3-think
1383
1334–1427
57
134
3
29%… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-britannica-nuextract3-sweep.bpl-ocr-bench-results-qwen35ocr-bench-rubenstein-judge-30bocr-bench-judge-eval-35bocr-bench-rubenstein-judge-kimi-k25robust-ocr-bench
