datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.CS2-HUD-OCR-Crops
CS2 HUD OCR Crops
Per-region HUD crops sliced from three Counter-Strike 2 match recordings,
labelled where possible from the demo file's parse_ticks state. Built
to train a specialist CRNN that replaces the EasyOCR killfeed reader
(currently ~6 s p95 on CPU) with a sub-30 ms specialist.
Source
Three matches by the same POV player (farouqqq), recorded in CS2's
built-in DVR + the corresponding .dem files:
sample
map
dem
rounds
resolution
fps
sample1
Ancient… See the full description on the dataset page: https://huggingface.co/datasets/ybashir/CS2-HUD-OCR-Crops.Burmese-Classics-OCR-RAW
Burmese (Myanmar) Books Dataset – Burmese Classics OCR (Daily Rolling Project)
Overview
A daily rolling dataset of Burmese books, built for AI, OCR, and NLP research.Each entry includes OCR text with metadata: title, author, page index, and Burmese character ratio.
This project fills a critical gap in Burmese-language resources:
Scarcity of public-domain Burmese text.
High technical and financial barriers to corpus building.
Enables incremental, open access for… See the full description on the dataset page: https://huggingface.co/datasets/minthanthtoo-cs/Burmese-Classics-OCR-RAW.persian-ocr-community-dataset-layout
Persian OCR Community Layout Annotations
Resumable layout annotations for the page images in
Reza2kn/persian-ocr-community-dataset.
Each row points to an exact source dataset revision, Parquet shard, blob, and row. It includes the
page identifier, page dimensions, handwriting flag, and structured layout boxes produced by
datalab-to/surya_layout2 at confidence threshold
0.4.
The boxes field contains label, confidence, raster-order position, and pixel coordinates
x0, y0, x1, y1.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-community-dataset-layout.ca_1860s_ocr_v1kat57-ocr-bench-500-results
Kat57 OCR benchmark — CER/WER
Strict reference-based evaluation of 16 OCR models on a deterministic 500-card sample from Lund University Library's Kat57 catalogue-card collection. This result set contains only Character Error Rate (CER) and Word Error Rate (WER); it does not contain VLM judging or ELO ratings.
The sample was drawn with seed 57 from tadad/kat57-ground-truth and is published as tadad/kat57-ground-truth-500. The OCR outputs are retained in… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-500-results.ca_1880s_ocr_v2MultiFinBen_OCR_Taskarmenian-ocr-crops
Tetrak Armenian OCR crops
Training data for tetrak_hy, the Armenian text recogniser we are
building as an EasyOCR custom model in
tetrak-hy-trainer
for Tetrak, an OCR pipeline for community
archives.
The dataset has three configurations:
corpus — 1,190 proofread pages of the Armenian Soviet
Encyclopedia, as plain text with full Wikisource provenance.
crops — the v0 synthetic pre-training set: 181,800 rendered
word crops with transcriptions.
crops-v1 — the v1 synthetic training… See the full description on the dataset page: https://huggingface.co/datasets/tetrak/armenian-ocr-crops.ca_1830s_ocr_v2arabic-turath-ocrca_1860s_ocr_v2old-nogay-turkish-ocr-corpus
Old Nogay Turkish OCR Corpus
This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis.
We are proud to release this as a contribution to Turkish NLP, Turkic-language NLP, historical NLP, and Turkish studies. The goal is not to pretend that a small OCR corpus is a polished benchmark. The goal is more… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/old-nogay-turkish-ocr-corpus.kat57-ocr-bench-results
Kat57 OCR smoke benchmark results
Exact ground-truth scoring for a 50-card Tesseract integration run over tadad/kat57-ground-truth-smoke. OCR outputs are published in the tesseract config of tadad/kat57-ocr-bench.
Model
CER
WER
Evaluated
Empty outputs
Error sentinels
Skipped references
Tesseract 5
0.4656
0.8605
50
1
0
0
The corpus totals are 4,629 character edits over 9,941 reference characters and 1,221 word edits over 1,419 reference words. Scoring used… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-results.English-PD-bad-OCRca_1870s_ocr_v214_Merged_RGB_ocrThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 126,
"total_frames": 54648,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:126"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/14_Merged_RGB_ocr.ca_1870s_ocr_v1pl-mixed-docs-ocr-dataset-100-v1-results
OCR Bench Results: Polish mixed documents benchmark
VLM-as-judge pairwise evaluation of OCR models on a small heterogeneous sample of Polish document-style images. Rankings depend strongly on document type, so this should be read as a document-specific OCR benchmark rather than a universal OCR ranking.
This benchmark uses a lightweight 100-image Polish OCR sample covering mixed document categories such as official forms, templates, certificates, structured layouts, invoices, and… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-mixed-docs-ocr-dataset-100-v1-results.Akkadian-OCR-Corpus
Akkadian OCR Corpus
OCR-extracted corpus from Old Assyrian cuneiform tablet publications and Akkadian linguistic resources, designed for LLM pretraining on ancient Mesopotamian languages.
Dataset Description
This dataset contains OCR-processed text from academic publications on Old Assyrian studies, including cuneiform tablet transliterations, translations, and the Chicago Assyrian Dictionary (CAD).
Splits
Split
Description
Pages
Characters
OCR Quality… See the full description on the dataset page: https://huggingface.co/datasets/lkevincc0/Akkadian-OCR-Corpus.ocr-bench-britannica-results-qwen35
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
zai-org/GLM-OCR
0.9B
1716
1673–1769
182
60
2
75%
2
lightonai/LightOnOCR-2-1B
1B
1697
1655–1749
158
61
1
72%
3
numind/NuExtract3
4B
1649
1604–1708
162
82
0
66%
4
FireRedTeam/FireRed-OCR
2.1B
1506
1469–1548
115
127
2
47%
5… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-britannica-results-qwen35.Aibot2_27Aug_STEA_Pick_OCRThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 10,
"total_frames": 3891,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/Aibot2_27Aug_STEA_Pick_OCR.OCR_BIMANUAL_22_21_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 89,
"total_frames": 34441,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:89"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/OCR_BIMANUAL_22_21_merged.ca_1890s_ocr_v2ocr-bench-judge-eval-27b
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
ELO
95% CI
Wins
Losses
Ties
Win%
1
lightonai/LightOnOCR-2-1B
1675
1571–1836
26
9
1
72%
2
FireRedTeam/FireRed-OCR
1612
1518–1767
25
13
1
64%
3
zai-org/GLM-OCR
1594
1480–1739
24
14
1
62%
4
deepseek-ai/DeepSeek-OCR
1437
1332–1546
15
23
1
38%
5
rednote-hilab/dots.ocr
1182… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-judge-eval-27b.14_RGB_OCRThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 62,
"total_frames": 25168,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:62"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/14_RGB_OCR.bpl-ocr-bench-results
OCR Bench Results: bpl-ocr-bench
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
ELO
95% CI
Wins
Losses
Ties
Win%
1
lightonai/LightOnOCR-2-1B
1559
1497–1630
39
25
0
61%
2
zai-org/GLM-OCR
1535
1471–1591
48
35
1
57%
3
rednote-hilab/dots.ocr
1453
1385–1515
26
37
0
41%
4
deepseek-ai/DeepSeek-OCR
1452
1388–1514
33
49
1
40%
Details
Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/bpl-ocr-bench-results.telugu-ocr-text-uncombinedocr-bench-britannica-results
OCR Bench Results: ocr-bench-britannica
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
rednote-hilab/dots.mocr
3B
1745
1714–1782
436
141
4
75%
2
lightonai/LightOnOCR-2-1B
1B
1741
1709–1779
426
141
4
75%
3
zai-org/GLM-OCR
0.9B
1738
1707–1773
469
157
2
75%
4
allenai/olmOCR-2-7B-1025-FP8
1719
1688–1753
454
167
4
73%… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ocr-bench-britannica-results.22_07_2026_OCR_BIMANUALThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 59,
"total_frames": 28575,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:59"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/22_07_2026_OCR_BIMANUAL.
