datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
permis-ocr-nuextract3
NuExtract3 on mehdigououiad/permis-ocr-bench
This dataset contains outputs from mehdigououiad/permis-ocr-bench processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: mehdigououiad/permis-ocr-bench
Model: numind/NuExtract3
Mode: markdown
Number of Samples: 2
Processing Time: 4.4 min
Processing Date: 2026-08-11 13:55 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mehdigououiad/permis-ocr-nuextract3.0004-pdf-nuextract3
NuExtract3 on stephenmcconnachie/0004-pdf-pages-test
This dataset contains outputs from stephenmcconnachie/0004-pdf-pages-test processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: stephenmcconnachie/0004-pdf-pages-test
Model: numind/NuExtract3
Mode: structured-extraction
Number of Samples: 50
Processing Time: 4.8 min
Processing Date: 2026-06-20 10:04 UTC
Configuration
Image Column:… See the full description on the dataset page: https://huggingface.co/datasets/stephenmcconnachie/0004-pdf-nuextract3.doab-nuextract3-smoke-output
NuExtract3 on davanstrien/doab-nuextract3-smoke-input
This dataset contains outputs from davanstrien/doab-nuextract3-smoke-input processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: davanstrien/doab-nuextract3-smoke-input
Model: numind/NuExtract3
Mode: structured-extraction
Number of Samples: 50
Processing Time: 5.3 min
Processing Date: 2026-05-21 09:35 UTC
Configuration
Image… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/doab-nuextract3-smoke-output.ocr-bench-britannica-nuextract3-sweep
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
glm-ocr
1556
1513–1599
116
78
0
60%
2
lighton-ocr-2
1538
1496–1580
99
76
0
57%
3
nuextract3
1532
1495–1573
104
82
7
54%
4
nuextract3-t0rep
1491
1451–1531
91
97
6
47%
5
nuextract3-think
1383
1334–1427
57
134
3
29%… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-britannica-nuextract3-sweep.nuextract3-extract-smoke
NuExtract3 on davanstrien/ufo-ColPali
This dataset contains outputs from davanstrien/ufo-ColPali processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: numind/NuExtract3
Mode: structured-extraction
Number of Samples: 5
Processing Time: 4.3 min
Processing Date: 2026-05-19 17:54 UTC
Configuration
Image Column: image
Output Column: extraction
Dataset Split: train… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nuextract3-extract-smoke.nuextract3-nls-cards-extract-smoke
NuExtract3 on NationalLibraryOfScotland/nls-index-cards-object-detection
This dataset contains outputs from NationalLibraryOfScotland/nls-index-cards-object-detection processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: NationalLibraryOfScotland/nls-index-cards-object-detection
Model: numind/NuExtract3
Mode: structured-extractionNumber of Samples: 5
Processing Time: 4.5 min
Processing Date: 2026-05-19… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nuextract3-nls-cards-extract-smoke.nuextract3-cards-verbose
NuExtract3 on davanstrien/nls-index-cards-only
This dataset contains outputs from davanstrien/nls-index-cards-only processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: davanstrien/nls-index-cards-only
Model: numind/NuExtract3
Mode: structured-extraction
Number of Samples: 49
Processing Time: 5.9 min
Processing Date: 2026-05-20 09:11 UTC
Configuration
Image Column: image
Output Column:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nuextract3-cards-verbose.nuextract3-cards-terse
NLS Advocates Library index cards → structured JSON (NuExtract3)
Demo output: scanned manuscript index cards from the National Library of Scotland's
Advocates Library, run through NuExtract3
(4B, Apache-2.0) for schema-guided structured extraction. Each row pairs the card
image with extraction — the JSON the model returned for that card.
How it was made
Input: NationalLibraryOfScotland/nls-index-cards-object-detection, filtered to the 49 pages that contain a card… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nuextract3-cards-terse.nuextract3-md-smoke
NuExtract3 on davanstrien/ufo-ColPali
This dataset contains outputs from davanstrien/ufo-ColPali processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: numind/NuExtract3
Mode: markdown
Number of Samples: 5
Processing Time: 4.5 min
Processing Date: 2026-05-19 17:44 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Temperature: 0.2… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nuextract3-md-smoke.humatheque-vlm-pred-nuextract3nuextract3-nls-bucket-smoke
NuExtract3 on NationalLibraryOfScotland/nls-index-cards-object-detection
This dataset contains outputs from NationalLibraryOfScotland/nls-index-cards-object-detection processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: NationalLibraryOfScotland/nls-index-cards-object-detection
Model: numind/NuExtract3
Mode: structured-extractionNumber of Samples: 5
Processing Time: 4.4 min
Processing Date: 2026-05-20… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nuextract3-nls-bucket-smoke.bpl-shelf-list-nuextract3
NuExtract3 on davanstrien/bpl-shelf-list-sample
This dataset contains outputs from davanstrien/bpl-shelf-list-sample processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: davanstrien/bpl-shelf-list-sample
Model: numind/NuExtract3
Mode: structured-extraction
Number of Samples: 120
Processing Time: 5.1 min
Processing Date: 2026-05-21 10:05 UTC
Configuration
Image Column: image
Output Column:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/bpl-shelf-list-nuextract3.nuextract3-nls-v2schema-smoke
NuExtract3 on NationalLibraryOfScotland/nls-index-cards-object-detection
This dataset contains outputs from NationalLibraryOfScotland/nls-index-cards-object-detection processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: NationalLibraryOfScotland/nls-index-cards-object-detection
Model: numind/NuExtract3
Mode: structured-extractionNumber of Samples: 5
Processing Time: 4.4 min
Processing Date: 2026-05-19… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nuextract3-nls-v2schema-smoke.NuExtract3.4_27B-SFT_predictionsdoab-nuextract3-smoke-input
