datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ocr-bench-britannica-nuextract3-sweep
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
glm-ocr
1556
1513–1599
116
78
0
60%
2
lighton-ocr-2
1538
1496–1580
99
76
0
57%
3
nuextract3
1532
1495–1573
104
82
7
54%
4
nuextract3-t0rep
1491
1451–1531
91
97
6
47%
5
nuextract3-think
1383
1334–1427
57
134
3
29%… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-britannica-nuextract3-sweep.NuExtract3.4_27B-SFT_predictions
