datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
post-ocr2OCRBenchGithub|Paper
OCRBench has been accepted by Science China Information Sciences.
short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.persian-ocr-community-dataset-argilla
Persian OCR community dataset - Argilla view
Lightweight two-column view for Argilla. image is an HF-hosted asset URL and label is JSON containing spatial OCR objects.
OCR-VQA
Dataset Card for "OCR-VQA"
More Information needed
pubmed-ocr
PubMed-OCR: PMC Open Access OCR Annotations
PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page is rendered to an image and annotated with Google Cloud Vision OCR, released in a compact JSON schema with word-, line-, and paragraph-level bounding boxes.
Scale (release):
209.5K articles
~1.5M pages
~1.3B words (OCR tokens)
This dataset is intended to support layout-aware modeling, coordinate-grounded QA, and… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/pubmed-ocr.LaTeX_OCR1% sampled from https://huggingface.co/datasets/linxy/LaTeX_OCR
ia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining.
@misc{moondream_ia_ocr,
author = {Vikhyat Korrapati},
title = {IA OCR Dataset},
year = {2025},
url = {https://huggingface.co/datasets/moondream/ia_ocr},
note = {Accessed: 2025-03-07}
}
persian-ocr-community-datasetanimetext-ocrOCRBench_v2blip3-ocr-200m
BLIP3-OCR-200M Dataset
Overview
The BLIP3-OCR-200M dataset is designed to address the limitations of current Vision-Language Models (VLMs) in processing and interpreting text-rich images, such as documents and charts. Traditional image-text datasets often struggle to capture nuanced textual information, which is crucial for tasks requiring complex text comprehension and reasoning.
Key Features
OCR Integration: The dataset incorporates Optical Character… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-ocr-200m.ocr_arabic_books
Arabic OCR Books Dataset (ocr_arabic_books)
This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions.
📚 Master Book Inventory
Total running pages in repository: 181,427
#
Book Name (English)
Book Name (Arabic)
Subset / Config Name
Page Count
Image Index Range
1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.sbb-dc-ocr
Dataset Card for Berlin State Library OCR data
Dataset Summary
The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945.
At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.
For each page with OCR text, the language has been determined by langid (Lui/Baldwin 2012).
Supported Tasks and Leaderboards
language-modeling: this dataset has the potential to be used… See the full description on the dataset page: https://huggingface.co/datasets/SBB/sbb-dc-ocr.OCR-Data
OCR Text Detection and Recognition Dataset
Dataset Description
A large-scale, multi-source OCR dataset aggregating 14 public benchmarks for text detection and recognition in both scene images and handwritten documents. Each image is paired with:
Transcribed text for each text region
Bounding boxes (axis-aligned rectangles) for each text region
Polygon coordinates (precise boundary points) for each text region
The dataset is stored in HuggingFace Parquet format with… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/OCR-Data.persian-printed-ocr-3.5m
Persian Printed OCR 3.5M
A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public
datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733
rejected rows are excluded. The viewer exposes exactly image and label.
Sources
AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0)
hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.OCR-liboaccn-OPUS-MIT-5M-clean
Description
This dataset is a processed version of liboaccn/OPUS-MIT-5M to make it easier to use, particularly for a visual question answering task where answer is an OCR transcription.Specifically, the original dataset has been processed to provide the image directly as a PIL rather than a path in an image column.We've also created a question column containing around 40 prompts based on via tutoiement, vouvoiement and imperative forms.
Note that this dataset contains only the… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/OCR-liboaccn-OPUS-MIT-5M-clean.bengali-ocr-synthetic
Bengali OCR Synthetic Dataset
A high-quality synthetic Bengali OCR dataset for fine-tuning vision-language models like DeepSeek-OCR 2. Generated using 100+ professional Bengali Unicode fonts and 13K+ unique Bengali words with advanced text rendering via FreeType and HarfBuzz.
Dataset Overview
Language: Bengali (বাংলা)
Task: Optical Character Recognition (OCR)
Format: Conversation-based (vision-language)
Total Samples: 30,000
Train: 27,007 samples
Validation: 2,993… See the full description on the dataset page: https://huggingface.co/datasets/rifathridoy/bengali-ocr-synthetic.ramanv-document-ocr-2LaTeX_OCR
LaTeX OCR 的数据仓库
本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。
如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++
后续追加新的数据也会放在这个仓库 ~~
原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR.
数据集
本仓库有 5 个数据集
small 是小数据集,样本数 110 条,用于测试
full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。
synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/linxy/LaTeX_OCR.deepseek-ocr-artifacts-test-XXOCRBench-v2parsynth-ocr-200kParsynthOCR is a synthetic dataset for Persian OCR. This version is a preview of the original 4 million samples dataset (ParsynthOCR-4M).
Usage
🤗 Datasets
from datasets import load_dataset
dataset = load_dataset("hezarai/parsynth-ocr-200k")
Hezar
pip install hezar
from hezar.data import Dataset
dataset = Dataset.load("hezarai/parsynth-ocr-200k", split="train")
ocr-document-processing-eval
ocr_document_processing_eval
Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks.
Repo: himalaya-ai/ocr-document-processing-eval
Task: document_processing_ocr
Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns.
Optional fine-tuning/eval file: *.sharegpt.json with messages and images.
Core Columns
id: unique sample identifier
image: relative path to the image file
ocr:… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/ocr-document-processing-eval.synthetic-receipts-ocr
synthetic-receipts-ocr
32,000 synthetic thermal receipts across 5 locales (US/UK/DE/IT/FR) — each a clean render plus a photo-degraded twin, with pixel-exact word boxes, full transcription, and structured KIE fields.
Samples train-000357 (US), train-000073 (UK), eval-001179 (DE), train-000222 (IT), train-000711 (FR) — real dataset rows, not mockups. Each receipt is its sample's image_photo, cut out along its own homography quad; no retouching beyond composition.
Built for… See the full description on the dataset page: https://huggingface.co/datasets/albertobarnabo/synthetic-receipts-ocr.myanmar-ocr-dataset-for-vlm
Myanmar OCR Dataset
A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts.
Subsets
Subset
Description
Details
single_font
Rendered with Pyidaungsu font only
437 books
multi_font
Rendered with 76 Myanmar fonts
3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.indic-deva-ocr-eval
indic_deva_eval
Broad Indic Devanagari OCR benchmark across printed pages, digits, word crops, and handwriting.
Repo: himalaya-ai/indic-deva-ocr-eval
Task: indic_devanagari_ocr
Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns.
Optional fine-tuning/eval file: *.sharegpt.json with messages and images.
Core Columns
id: unique sample identifier
image: relative path to the image file
ocr: ground-truth text label… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/indic-deva-ocr-eval.Llama-Nemotron-VLM-Dataset-v1-OCR4al-kawakib-magazine-ocr
Al-Kawakib Magazine OCR Pages
This dataset contains rendered Arabic magazine page images paired with page-level text and line-level bounding boxes. It is intended for OCR, document understanding, and VLM fine-tuning experiments.
Fine-Tuning Notebook
A standalone Google Colab notebook for DeepSeek-OCR 3B + TRL SFT is available here:
Open the fine-tuning notebook in Colab
The notebook can train on this dataset alone or on all three Arabic magazine OCR… See the full description on the dataset page: https://huggingface.co/datasets/amrosama/al-kawakib-magazine-ocr.07-08-2026_OCR_BimanualThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 37,
"total_frames": 17770,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:37"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/07-08-2026_OCR_Bimanual.
